Trajectory prediction method, model training method, device and equipment for traffic participants
By directly obtaining perceptual vectors from the backbone network of the perceptual model in the autonomous driving system, preprocessing steps are avoided, and the accuracy and real-time nature of trajectory prediction are improved, and the accuracy reduction problem caused by information loss in the prior art is solved.
Patent Information
- Application Number
- CN202510341273.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the trajectory prediction module in the autonomous driving system loses information due to preprocessing operations, which reduces the accuracy of trajectory prediction.
By directly obtaining perceptual vectors from the backbone network of the perceptual model, avoiding complex preprocessing steps, directly performing prediction inference, and obtaining the predictive trajectory of traffic participants.
It improves the accuracy and real-time trajectory prediction and enhances the response ability of autonomous driving systems in complex traffic environments.
Smart Images

Figure CN120299001A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and particularly to a method for predicting the trajectory of traffic participants, a model training method, a device, and a device. Background Art
[0002] Trajectory prediction plays a key role in an autonomous driving system and is a bridge connecting environmental perception and decision-making and planning. Through accurate trajectory prediction, an autonomous driving vehicle can better cope with complex traffic environments and improve driving safety and efficiency. To achieve this goal, an autonomous driving system usually adopts a modular design, including independent modules such as perception, prediction, planning, and control. The perception module is responsible for processing data obtained by perception devices (such as cameras, lidars, and millimeter-wave radars) to extract structured data of traffic participants. Next, the prediction module uses this structured data to predict the trajectories of traffic participants, providing a basis for subsequent planning and control.
[0003] In related technologies, in order to ensure that the prediction module can effectively process structured data, it is usually necessary to perform preprocessing operations such as normalizing and feature encoding on the feature data output by the perception module. However, the preprocessing operations will cause loss of information, thereby reducing the accuracy of trajectory prediction. Summary of the Invention
[0004] This application provides a method for predicting the trajectory of traffic participants, a model training method, a device, and a device, so as to improve the accuracy of predicting the trajectory of traffic participants and reduce the time consumption of trajectory prediction.
[0005] In a first aspect, this application provides a method for predicting the trajectory of traffic participants, which is applied to a vehicle. The method includes:
[0006] Obtain perception data, where the perception data includes image data and / or point cloud data of traffic participants;
[0007] Infer the perception data through the backbone network of the perception model to obtain a perception vector; the perception vector is used to represent the characteristics of the traffic participants and the traffic environment;
[0008] Perform prediction inference on the perception vector to obtain the predicted trajectory of the traffic participant.
[0009] In a possible implementation manner, the performing prediction inference on the perception vector to obtain the predicted trajectory of the traffic participant includes:
[0010] Perform prediction inference on the perception vector according to the trajectory prediction model to obtain the predicted trajectory of the traffic participant;
[0011] Among them, the trajectory prediction model is a model trained from a deep learning model for trajectory prediction based on a perception vector.
[0012] In a possible implementation, processing the perception data through a backbone network of a perception model to obtain a perception vector includes:
[0013] Determine a target layer in the backbone network of the perception model, where the target layer is a feature processing layer with the richest semantic features output in the backbone network;
[0014] Perform forward propagation calculation on the perception data to obtain an intermediate vector corresponding to the perception data;
[0015] Infer the intermediate vector through the target layer to obtain the perception vector.
[0016] In a possible implementation, the feature processing layer is a convolutional layer or a fully connected layer.
[0017] In a possible implementation, the method further includes:
[0018] Receive a trajectory prediction model sent by a server and deploy the trajectory prediction model.
[0019] In a second aspect, the present application provides a model training method applied to a server, and the method includes:
[0020] Obtain a plurality of training samples, where each training sample includes a perception vector and a driving trajectory corresponding to the perception vector, and the perception vector is used to represent the characteristics of traffic participants and the traffic environment;
[0021] Train a deep learning model according to the plurality of training samples to obtain a trajectory prediction model.
[0022] In a possible implementation, the method further includes:
[0023] Send the trajectory prediction model to at least one vehicle.
[0024] In a third aspect, the present application provides a trajectory prediction device for traffic participants, and the device includes:
[0025] A first processing module configured to obtain perception data, where the perception data includes image data and / or point cloud data of traffic participants;
[0026] A second processing module configured to infer the perception data through a backbone network of a perception model to obtain a perception vector; the perception vector is used to represent the characteristics of the traffic participants and the traffic environment;
[0027] A prediction module, configured to perform prediction inference on the perception vector to obtain the predicted trajectory of the traffic participant.
[0028] In a fourth aspect, the present application provides a model training device, which includes:
[0029] An acquisition module, configured to acquire a plurality of training samples, each of which includes a perception vector and the driving trajectory corresponding to the perception vector, and the perception vector is used to represent the characteristics of traffic participants and the traffic environment;
[0030] A processing module, configured to train a deep learning model according to the plurality of training samples to obtain a trajectory prediction model.
[0031] In a fifth aspect, the present application provides a vehicle, which includes: a processor, a memory, and a communication interface;
[0032] The memory stores computer-executable instructions;
[0033] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspect.
[0034] In a sixth aspect, the present application provides a server, which includes: a processor, a memory, and a communication interface;
[0035] The memory stores computer-executable instructions;
[0036] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the second aspect.
[0037] In a seventh aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of the first aspect or the second aspect.
[0038] In an eighth aspect, the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it is used to implement the method according to any one of the first aspect or the second aspect.
[0039] The trajectory prediction method, model training method, device, and equipment for traffic participants provided by this application can obtain perception data, which may include image data and / or point cloud data of traffic participants. The perception data is inferred through the backbone network of the perception model to obtain a perception vector. Furthermore, the perception vector is inferred to obtain the predicted trajectory of the traffic participant. By directly extracting the perception vector from the backbone network, information loss caused by preprocessing is effectively avoided, thereby improving the accuracy of trajectory prediction. At the same time, directly using the perception vector as the input of the prediction model avoids the complex preprocessing process of encoding the structured information output by perception into the prediction input vector, effectively reducing the time consumption of the prediction module. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0041] Figure 1 Schematic diagram of related technologies;
[0042] Figure 2 Schematic flowchart of the first embodiment of the trajectory prediction method for traffic participants provided by this application;
[0043] Figure 3 Schematic flowchart of the second embodiment of the trajectory prediction method for traffic participants provided by this application;
[0044] Figure 4 Schematic diagram of the structure of a convolutional neural network provided by an embodiment of this application;
[0045] Figure 5 Schematic flowchart of the model training method provided by an embodiment of this application;
[0046] Figure 6 Schematic diagram of the framework for predicting the trajectory of traffic participants provided by an embodiment of this application;
[0047] Figure 7 Schematic diagram of the structure of the trajectory prediction device for traffic participants provided by an embodiment of this application;
[0048] Figure 8 Schematic diagram of the structure of the model training device provided by an embodiment of this application;
[0049] Figure 9 Schematic diagram of the structure of the electronic device provided by this application.
[0050] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0051] Here, exemplary embodiments will be described in detail, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0052] The intelligent development of autonomous driving systems depends to a large extent on the understanding of the surrounding environment and the optimization of vehicle behavior decisions. To achieve this, the trajectory prediction of traffic participants becomes a key link because it directly affects the vehicle's response ability in complex traffic environments.
[0053] Modern autonomous driving systems usually adopt a modular technical framework, which means that the system is divided into multiple independent parts, and each part is responsible for a specific function. Among these modules, the perception module is responsible for collecting and fusing data through various perception devices (such as lidar, millimeter-wave radar, and cameras), and then outputting feature data such as the position, size, orientation angle, and speed of each traffic participant. Next, the prediction module converts the structured data provided by the perception module into an analysis of the future movement of traffic participants, so as to provide a basis for the motion planning of the host vehicle.
[0054] In the related art, in order to ensure that the prediction module can effectively process the feature data, preprocessing operations such as data standardization and feature vectorization need to be performed on the structured data (the position, size, orientation angle, speed, etc. of traffic participants) output by the original perception. For the sake of understanding, hereinafter, taking obstacles as traffic participants as an example, the solution for obstacle trajectory prediction in the related art will be described in detail.
[0055] Figure 1 For the related art schematic diagram. Please refer to Figure 1 , the perception module (also the perception model) can receive image data and point cloud data, and obtain perception information such as obstacles and lane lines after reasoning. Before the prediction module (also the prediction model) reasons about the perception information, operations such as normalization and feature encoding need to be performed on the perception information to ensure that the prediction module can effectively perform subsequent processing.
[0056] However, the above data conversion process is prone to cause attenuation of kinematic features (such as smoothing of speed fluctuation characteristics) and dilution of scene interaction information (such as discretized representation of vehicle-pedestrian game intentions), thereby reducing the detailedness of the dynamic behavior prediction of traffic participants, ultimately resulting in a deviation between the predicted trajectory and the true trajectory of traffic participants, and thus reducing the accuracy of trajectory prediction.
[0057] To address the above problems, the inventors considered that perceptual vectors for representing traffic participants and their environmental features could be directly obtained from the backbone network of the perception model. Then, direct prediction inference is performed on these perceptual vectors to better capture the dynamic behavior of traffic participants and scene interaction information, thereby improving the accuracy of trajectory prediction. Accordingly, after multiple experiments, the inventors found that after obtaining perceptual data including image data and / or point cloud data of traffic participants, the perceptual data can be inferred through the backbone network of the perception model to obtain perceptual vectors, and then prediction inference is performed on the perceptual vectors to obtain the predicted trajectory of traffic participants, thereby effectively reducing information loss during the data conversion process and improving the accuracy of the trajectory prediction of traffic participants. Based on this, the present application proposes a method for predicting the trajectory of traffic participants, aiming to enhance the prediction ability of traffic participant trajectories by optimizing the perception and prediction inference steps.
[0058] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0059] Figure 2 It is a schematic flowchart of the first embodiment of the method for predicting the trajectory of traffic participants provided by the present application. Please refer to Figure 2 , the method may include:
[0060] S201. Obtain perceptual data, where the perceptual data includes image data and / or point cloud data of traffic participants.
[0061] The execution subject of the embodiments of the present application can be a vehicle or a device for predicting the trajectory of traffic participants disposed in the vehicle. The device for predicting the trajectory of traffic participants can be implemented by software or a combination of software and hardware. The device for predicting the trajectory of traffic participants can be a processor in the vehicle. For the convenience of understanding, in the following, the execution subject is taken as an example of a vehicle for description.
[0062] In this step, perceptual data can be obtained according to the perceptual devices disposed on the vehicle, and the perceptual data can include image data and / or point cloud data of traffic participants.
[0063] In autonomous driving, perception devices are one of the core components, responsible for collecting environmental information and supporting the vehicle's decision-making process. These devices typically include cameras, lidars, radars, and ultrasonic sensors, which can generate various types of data, such as image data and point cloud data, to help the vehicle understand its surrounding environment. Image data is mainly captured by cameras, which can be monocular, stereo, or surround-view systems. Image data is used to identify and classify traffic participants, such as road signs, traffic lights, pedestrians, vehicles, and other obstacles, and can also be used for lane detection and object tracking.
[0064] On the other hand, point cloud data is mainly generated by lidars. Lidars construct a three-dimensional point cloud by emitting laser beams and measuring their return times, providing high-precision three-dimensional environmental modeling to help identify the shape, size, and position of traffic participants.
[0065] For ease of understanding, an example where the perception data includes image data is used for illustration.
[0066] In an alternative embodiment, the perception data may include consecutive frames of image data. By analyzing these consecutive frames of image data, the system can gain a deeper understanding of the dynamic changes and object motion characteristics in the environment. By correlating the positions of traffic participants in multiple frames, the system can continuously track the motion paths of objects. This is crucial for predicting the future positions of objects, especially in complex traffic environments, and can help the vehicle make safer driving decisions.
[0067] The number of consecutive frames of image data can be preset according to user requirements and application scenarios to optimize the performance and response capabilities of the autonomous driving system. For example, in a highway driving scenario, since the vehicle is traveling at a high speed, the system may require a higher frame rate to capture the rapidly changing environment, so it can be preset to 20 frames per second or higher. This high frame rate can provide finer temporal resolution and help the system quickly identify and respond to surrounding dynamic changes, such as a vehicle approaching rapidly or an obstacle suddenly appearing.
[0068] For example, the vehicle can obtain perception data through a camera installed on the vehicle. The perception data includes 20 consecutive frames of image data, and the image data may include pedestrian A.
[0069] S202. Infer the perception data through the backbone network of the perception model to obtain a perception vector.
[0070] In this step, the perception data can be inferred through the backbone network of the perception model to obtain a perception vector representing the characteristics of traffic participants and the traffic environment.
[0071] Optionally, the perception vector is usually a multi-dimensional vector, with each dimension representing a specific feature. These vectors can be of a fixed length, facilitating subsequent processing and analysis.
[0072] The backbone network of the perception model is a core component in the deep learning model, specifically designed to extract features from the input data. It typically consists of a series of hierarchical feature processing layers that can gradually abstract and compress the original information of the input data, generating high-level feature representations for subsequent task processing.
[0073] In an alternative implementation, the backbone network usually consists of convolutional layers, pooling layers, and fully connected layers, which are organized in a hierarchical manner to gradually extract the spatial and semantic features of the data.
[0074] For example, inference can be performed on 20 consecutive frames of image data through the convolutional layer in the backbone network of the perception model to obtain the perception vector. In a specific implementation, inference can be performed on 20 frames of image data at once through the convolutional layer in the backbone network of the perception model to obtain the perception vector; in addition, inference can also be performed on each frame of image data through the convolutional layer in the backbone network of the perception model to obtain a single-frame vector corresponding to a single frame of the image, and the single-frame vectors corresponding to the 20 frames of image data are concatenated to obtain the final perception vector.
[0075] S203. Perform predictive inference on the perception vector to obtain the predicted trajectory of the traffic participant.
[0076] In this step, inference can be performed on the perception vector to obtain the predicted trajectory of the traffic participant. In a specific implementation, predictive inference can be performed on the perception vector according to the trajectory prediction model to obtain the predicted trajectory of the traffic participant. Among them, the trajectory prediction model is a model for trajectory prediction based on the perception vector obtained by training the deep learning model.
[0077] For example, the predicted trajectory of pedestrian A can be obtained by performing predictive inference on the perception vector according to the trajectory prediction model.
[0078] In an alternative implementation, the vehicle can plan its own driving trajectory based on the predicted trajectory of the traffic participant obtained by the trajectory prediction model, thereby improving driving safety and efficiency and enhancing the vehicle's response ability in the dynamic traffic environment.
[0079] In the embodiments of the present application, a vehicle can infer the acquired perception data through the backbone network of a perception model to obtain a perception vector. The perception data includes image data and / or point cloud data of traffic participants. Further, the perception vector can be predicted and inferred to generate a predicted trajectory of the traffic participant. In the above process, the vehicle directly obtains the perception vector from the backbone network of the perception model, avoiding complex pre-processing before prediction, effectively reducing information loss, and improving the accuracy of trajectory prediction.
[0080] Based on the Figure 2 embodiment shown below, in combination with Figure 3 , the above method for predicting the trajectory of a traffic participant will be further described in detail.
[0081] Figure 3 FIG. Figure 3 is a schematic flowchart of the second embodiment of the method for predicting the trajectory of a traffic participant provided by the present application. Please refer to
[0082] S301. Receive a trajectory prediction model sent by a server and deploy the trajectory prediction model.
[0083] In this step, the vehicle can receive the trajectory prediction model sent by the server and deploy it in the vehicle's autonomous driving system.
[0084] S302. Acquire perception data, where the perception data includes image data and / or point cloud data of traffic participants.
[0085] Optionally, the parameters of the cache module can be modified in the vehicle's operating environment, and the perception data can be acquired by caching the image data and / or point cloud data of traffic participants.
[0086] For example, the cache module can cache perception data, and the perception data includes image data of pedestrian A and vehicle B.
[0087] S303. Determine a target layer in the backbone network of the perception model, where the target layer is the feature processing layer that outputs the richest semantic features in the backbone network.
[0088] In this step, the feature processing layer that outputs the richest semantic features in the backbone network of the perception model can be selected as the target layer. Among them, the feature processing layer can be a convolutional layer or a fully connected layer.
[0089] The feature processing layer with the richest output semantic features refers to the intermediate layer that can effectively capture high-level semantic information. This layer usually has the strongest semantic information extraction ability and can effectively capture complex high-level features, such as the shape, category, and complex patterns of objects, rather than just simple edges or textures. Optionally, forward propagation can be performed on each layer of the perception model to extract the output feature maps of each layer, and the semantic richness can be evaluated by calculating the entropy or information gain of these feature maps. By selecting the layer with the highest entropy or information gain, it can be determined which layer most effectively captures the complex semantic information of the perception data, and thus it can be used as the target layer.
[0090] For example, in the backbone network of the perception model, there are convolutional layer A, convolutional layer B, and convolutional layer C. Among them, if convolutional layer A is the feature processing layer with the richest output semantic features, then convolutional layer A can be selected as the target layer.
[0091] In another alternative implementation, if there are multiple feature processing layers in the backbone network of the perception model, the feature processing layer with the most output features can also be selected as the target layer among the multiple feature processing layers.
[0092] For example, in the backbone network of the perception model, there are 3 feature processing layers, namely convolutional layer A, convolutional layer B, and convolutional layer C. The output features of convolutional layer A are 30, the output features of convolutional layer B are 20, and the output features of convolutional layer C are 10; then convolutional layer A can be selected as the target layer.
[0093] In an alternative implementation, when selecting the target layer in the backbone network, a reasonable selection can also be made according to the specific task and network structure. The network structure can include an input layer, a backbone network, and an output layer, where the backbone network can include at least one of a convolutional layer, a pooling layer, an activation function layer, and a fully connected layer.
[0094] When the task requires extracting multi-level spatial features from the input data, the convolutional layer is the first choice. For example, in image classification, object detection, or semantic segmentation tasks, the convolutional layer can capture features such as edges, textures, and shapes through local perception. Shallow convolutional layers are suitable for extracting low-level features, while deep convolutional layers are suitable for extracting high-level abstract features. Therefore, if the task needs to focus on the local details or global structure of the input data, the convolutional layer can be selected as the target layer.
[0095] When the task requires reducing the spatial dimension of the feature map to reduce the computational amount and at the same time enhancing the robustness of the model to transformations such as translation and scaling, the pooling layer can be selected. The pooling layer reduces the size of the feature map through downsampling operations (such as max pooling or average pooling), while retaining important features. In image classification or object detection tasks, the pooling layer is usually used to gradually compress the spatial information of the feature map, thereby extracting more global features.
[0096] When the task requires introducing non-linear characteristics to enhance the model's expressive power, the activation function layer can be selected. The activation function layer performs non-linear transformation on the output of the convolutional layer or the fully connected layer, enabling the network to learn and represent complex patterns. For example, when capturing non-linear relationships in data (such as image classification, speech recognition), the activation function layer (such as ReLU, Sigmoid, or Tanh) is essential. The activation function layer closer to the input end can enhance the expressive power of low-level features, while the activation function layer closer to the output end helps in modeling high-level features.
[0097] When the task requires global integration of the features extracted by the previous layers and outputting the final classification or regression results, the fully connected layer can be selected. The fully connected layer flattens the feature map or feature vector into a one-dimensional vector and performs linear transformation through the weight matrix and bias vector, finally outputting the results required by the task. For example, in the image classification task, the fully connected layer is usually located at the end of the network, mapping high-level features to specific class labels. If the task requires comprehensive judgment of global features, the fully connected layer is an ideal choice.
[0098] S304. Perform forward propagation calculation on the perceptual data to obtain the intermediate vector corresponding to the perceptual data.
[0099] In this step, the forward propagation calculation can be performed on the perceptual data through the perceptual model to obtain the intermediate vector to be input into the target layer.
[0100] Figure 4 This application provides a schematic structural diagram of a convolutional neural network. Please refer to Figure 4 , including an input layer, a convolutional layer 1, a convolutional layer 2, a fully connected layer, and an output layer. Among them, the backbone network includes the convolutional layer 1, the convolutional layer 2, and the fully connected layer. The input layer is used to input perceptual data. The convolutional layer is used to perform convolution operations using convolutional kernels to extract corresponding image feature information. The fully connected layer is used to summarize the image feature information obtained by the convolutional layer. The output layer is used to output structured data based on the image feature information of the fully connected layer.
[0101] If the fully connected layer is selected as the target layer in the backbone network, where the fully connected layer is the feature processing layer with the richest output semantic features or the feature processing layer with the most output features. Then the outputs of the convolutional layer 1 and the convolutional layer 2 can be used as the intermediate vector. If the convolutional layer 1 is selected as the target layer in the backbone network, then the output of the input layer can be used as the intermediate vector. It should be understood that in a neural network, the main role of the input layer is to receive external data and pass it to the subsequent layers of the network. The input layer itself usually does not perform any calculations or transformations. It is just a data entry. Therefore, the "output" of the input layer is actually the input data itself, usually after some preprocessing to suit the needs of the network.
[0102] For example, according to Figure 4 the result of the convolutional neural network shown, if the fully connected layer is the target layer in the backbone network, then the input layer and the convolutional layer can be used to perform forward propagation calculations on the perception data in sequence, and the intermediate vector can be determined according to the output result of the convolutional layer.
[0103] S305. Infer the intermediate vector through the target layer to obtain the perception vector.
[0104] In this step, the intermediate vector can be inferred according to the target layer determined in the backbone network to obtain the perception vector, and the perception vector is used to represent the characteristics of traffic participants and the traffic environment.
[0105] For example, as Figure 4 shown, if the target layer is the fully connected layer in the backbone network, then the fully connected layer can be used to infer the intermediate vector to obtain the perception vector for representing pedestrian A and vehicle B.
[0106] S306. According to the trajectory prediction model, perform prediction inference on the perception vector to obtain the predicted trajectory of the traffic participant.
[0107] In this step, the vehicle can perform prediction inference on the perception vector determined through the intermediate layer according to the pre-deployed trajectory prediction model to obtain the predicted trajectory of the traffic participant.
[0108] For example, the vehicle can perform prediction inference on the perception vector determined through the target layer according to the pre-deployed trajectory prediction model to obtain the predicted trajectory 1 of pedestrian A and the predicted trajectory 2 of vehicle B.
[0109] In the embodiment of the present application, the vehicle can receive the trajectory prediction model sent by the server and deploy the trajectory prediction model. Further, it can obtain the perception data including the image data and / or point cloud data of the traffic participant, determine the target layer in the feature processing layer with the most output features in the backbone network of the perception model, infer the intermediate vector obtained from the perception data based on the target layer to obtain the perception vector; according to the trajectory prediction model, perform prediction inference on the perception vector to obtain the predicted trajectory of the traffic participant. In the above process, by directly selecting the target layer with the richest output semantic features in the backbone network of the perception model to infer the intermediate vector, generating the perception vector and using it for trajectory prediction, complex pre-processing steps before prediction are avoided. This method reduces information loss, simplifies the calculation process, improves the accuracy of trajectory prediction, reduces the time required for pre-processing before prediction, improves real-time performance, enhances the robustness of the model in complex traffic environments, and provides a more efficient and reliable solution for the autonomous driving system.
[0110] Figure 5The flowchart of the model training method provided by the embodiment of the present application. Please refer to Figure 5 , this method is applied to a server and includes:
[0111] S501. Obtain a plurality of training samples, where each training sample includes a perception vector and a driving trajectory corresponding to the perception vector.
[0112] In this step, the server can obtain a plurality of training samples, and each training sample includes a perception vector and its corresponding driving trajectory. The perception vector is used to represent an abstract representation of traffic participants (such as vehicles, pedestrians, bicycles, etc.) and traffic environments (such as road structures, traffic signals, weather conditions, etc.).
[0113] Optionally, the perception vector can be a feature representation obtained after reasoning on perception data collected by perception devices (such as cameras, lidars, radars, etc.). It can contain the following information:
[0114] Traffic participant features: such as the speed, acceleration, direction, size, type (sedan, truck, bicycle, etc.) of a vehicle, and the position, speed, posture, etc. of a pedestrian.
[0115] Traffic environment features: such as lane lines, traffic signal states, road types (highways, urban roads, etc.), weather conditions (sunny, rainy, foggy, etc.), obstacle positions, etc.
[0116] Time and space information: such as timestamps, geographical locations, relative distances, etc.
[0117] The perception vector is usually a high-dimensional vector and may include hundreds or even thousands of feature values. These feature values can be extracted and compressed by a perception model (such as a convolutional neural network, a recurrent neural network, etc.) for subsequent processing and analysis.
[0118] In an alternative implementation, a perception model is deployed in a vehicle. The vehicle can obtain perception data according to pre-set perception devices, determine a target layer in the backbone network of the perception model, reason on the perception data through the target layer to obtain a perception vector, and then upload the perception vector and its corresponding driving trajectory to the server.
[0119] In another alternative implementation, a perception model is deployed in the server. The vehicle can obtain perception data according to pre-set perception devices, upload the perception data and its corresponding driving trajectory to the server, and the server can reason on the perception data through the backbone network of the perception model set therein to obtain the perception vector corresponding to the perception data.
[0120] Optionally, for the driving trajectory, it can be the movement path of a traffic participant within a specific time period. It is usually composed of a series of time-position points, and each point includes information such as timestamp, longitude and latitude coordinates, speed, direction, etc.
[0121] In a specific implementation, the perception data can be inferred according to the perception model to obtain a perception vector as the input for training the trajectory prediction model. Among them, the perception data can include image data and / or point cloud data of consecutive frames; the future frame positions of the traffic participants (target objects) output by the perception model are spliced into a driving trajectory as the label for training the deep learning model. Furthermore, a dataset for training the trajectory prediction model can be determined according to multiple perception vectors and the driving trajectory corresponding to each perception vector. Among them, the driving trajectory spliced from the future frame positions includes part of the consecutive frames selected from the structured data output from all frames after the perception model infers the perception data.
[0122] For example, if the perception data contains 50 consecutive frames of image data, the perception model can be used to infer the first 20 frames of image data to extract the perception vector. Then, use the perception model to infer the last 30 frames of image data to predict the future positions of the traffic participants in these 30 frames, and splice these position data into a complete driving trajectory. This trajectory reflects the movement path of the traffic participant within the last 30 frames. Finally, take the perception vector obtained from the first 20 frames of image data as the input, and the driving trajectory generated from the last 30 frames as the label to form a training sample.
[0123] It should be noted that the perception model can perform one inference on the perception data containing 50 consecutive frames of image data to obtain the perception vector and its corresponding driving trajectory; in addition, in practical applications, the collection of perception data is often continuous, so more training samples can be generated by using a sliding window. By sliding a 50-frame window in the time series, a corresponding perception vector and structured data can be generated for each frame of data. For example, if the perception data includes 100 consecutive frames of image data, the first window ranges from frame 1 to frame 50, and the first training sample can be generated; slide the window by one frame, the second window ranges from frame 2 to frame 51, and the second training sample can be generated; continue to slide the window until the last window ranges from frame 51 to frame 100, and the 51st training sample can be generated.
[0124] S502. Train the deep learning model according to multiple training samples to obtain a trajectory prediction model.
[0125] In this step, the server can use numerous training samples to systematically train the deep learning model, thereby constructing an accurate trajectory prediction model.
[0126] Optionally, a suitable deep learning model architecture can be selected, such as a convolutional neural network, a recurrent neural network, a long short-term memory network, or a Transformer, etc. These models can effectively capture the spatio-temporal dependencies in time series data. During the training process, optimization algorithms and loss functions are used to iteratively optimize the model, gradually adjusting the model parameters to enable it to more accurately predict future trajectories. In addition, to prevent overfitting, regularization techniques or early stopping strategies can be adopted. Finally, through multiple trainings and validations, the model can exhibit good generalization ability on the test set, thus achieving accurate prediction of future trajectories.
[0127] Optionally, the optimization algorithm can be Adaptive Moment Estimation (Adam) or Stochastic Gradient Descent (SGD). The Adam algorithm can adjust the learning rate of each parameter by calculating the first-order moment estimate and the second-order moment estimate of the gradient. This method not only considers the average value of the gradient (first-order moment), but also considers the uncentered variance of the gradient (second-order moment), thus achieving an adaptive adjustment of the learning rate; SGD is an iterative optimization algorithm that randomly selects a sample in each iteration to calculate the gradient and update the model parameters. Compared with the traditional gradient descent algorithm, SGD does not need to calculate the gradients of all samples in each iteration, so the computational amount can be greatly reduced.
[0128] Optionally, the loss function can be mean squared error or cross entropy; the regularization technique can be L2 regularization or Dropout.
[0129] In a specific implementation manner, after the server finishes training the trajectory prediction model, it can send the trained trajectory prediction model to at least one vehicle so that the vehicle can locally and real-time apply the model for trajectory prediction.
[0130] In the embodiments of this application, the server can obtain multiple training samples and use these training samples to train a deep learning model, thereby obtaining a trajectory prediction model for predicting the trajectories of traffic participants. In the above process, by extracting the perception vectors from the backbone network, multi-level and multi-scale feature information can be obtained. This information can help the trajectory prediction model better understand the dynamic behaviors and environmental changes of traffic participants, thereby improving the accuracy and robustness of the prediction.
[0131] Furthermore, by deploying the model on the vehicle, the vehicle can reduce its dependence on the server, reduce communication latency, and continue to provide prediction functions in the case of poor or interrupted network connections. In addition, the vehicle can selectively update or replace the model according to its own computing power and storage resources to ensure that it always uses the latest and most effective prediction technologies.
[0132] Figure 6 Schematic diagram of the framework for predicting the trajectories of traffic participants provided by the embodiments of this application. Please refer to Figure 6 , in this framework, there are two schemes for predicting the trajectories of traffic participants, namely the scheme for predicting the trajectories of traffic participants provided by the related art and the scheme for predicting the trajectories of traffic participants provided by this application.
[0133] In the related art, the perception model can infer image data and / or point cloud data to extract the feature data of traffic participants (such as obstacle 1, obstacle 2, lane line 1, lane line 2). Subsequently, these feature data go through pre-prediction processing steps such as normalization, obstacle feature encoding, and lane line feature encoding to generate feature vectors. Furthermore, the prediction model can infer the feature vectors to obtain the predicted trajectories of traffic participants. In the related art, since pre-prediction processing operations such as normalization and feature encoding of the relevant features of traffic participants are required before the prediction model infers the image data and / or point cloud data, information loss will occur, thus reducing the accuracy of trajectory prediction.
[0134] The technical solution provided by this application includes: the training process of the deep learning model by the server, and the process of the vehicle predicting the trajectories of traffic participants according to the trained deep learning model.
[0135] During the training process, the perception vector can be extracted from the perception model through the data acquisition module, and the future trajectories of traffic participants can be obtained through the perception output of consecutive frames, and the training data can be constructed based on this to train the deep learning model to obtain the trajectory prediction model. Further, the server can send the trajectory prediction model to at least one vehicle, and after receiving the trajectory prediction model, the vehicle can deploy it on the vehicle side. Among them, the future trajectories of traffic participants (which can also be the trajectory ground truth for training the deep learning model) included in the training data can be obtained from the output results of the perception model.
[0136] During the process of predicting the trajectories of traffic participants, a vehicle can obtain image data and / or point cloud data of traffic participants, infer the perception data including the image data and / or point cloud data through the backbone network of a perception model to obtain a perception vector, and perform predictive inference on the perception vector according to a pre-deployed trajectory prediction model to obtain the predicted trajectories of traffic participants. Since the perception vector can be directly obtained from the backbone network without complex pre-prediction processing such as normalization, obstacle feature encoding, and lane line feature encoding, information loss can be effectively reduced, thereby improving the accuracy of predicting the trajectories of traffic participants.
[0137] The trajectory prediction framework for traffic participants provided by the embodiments of this application simplifies the data processing flow by directly obtaining the perception vector from the backbone network of the perception model, avoiding the complex pre-prediction processing steps in traditional methods. This method of directly obtaining the perception vector reduces information loss and improves the real-time performance and accuracy of trajectory prediction. In this way, the system can respond more quickly to changes in the dynamic traffic environment, improving the decision-making ability and safety of autonomous vehicles in complex traffic scenarios.
[0138] Figure 7 It is a schematic structural diagram of the trajectory prediction device for traffic participants provided by the embodiments of this application. Please refer to Figure 7 , the trajectory prediction device 10 for traffic participants may include:
[0139] A first processing module 11, configured to obtain perception data, where the perception data includes image data and / or point cloud data of traffic participants;
[0140] A second processing module 12, configured to infer the perception data through the backbone network of the perception model to obtain a perception vector; the perception vector is used to represent the characteristics of the traffic participant and the traffic environment;
[0141] A prediction module 13, configured to perform predictive inference on the perception vector to obtain the predicted trajectories of the traffic participants.
[0142] The trajectory prediction device for traffic participants provided by the embodiments of this application can execute the technical solutions shown in the above method embodiments, and its implementation principles and beneficial effects are similar, which will not be elaborated here.
[0143] In a possible implementation manner, the prediction module 13 is specifically configured to:
[0144] Perform predictive inference on the perception vector according to the trajectory prediction model to obtain the predicted trajectories of the traffic participants;
[0145] Among them, the trajectory prediction model is a model trained from a deep learning model for trajectory prediction based on perception vectors.
[0146] In a possible implementation manner, the second processing module 12 is specifically configured to:
[0147] Determine a target layer in the backbone network of the perception model, where the target layer is the feature processing layer with the richest output semantic features in the backbone network;
[0148] Perform forward propagation calculation on the perception data to obtain an intermediate vector corresponding to the perception data;
[0149] Infer the perception vector through the target layer for the intermediate vector.
[0150] In a possible implementation manner, the feature processing layer is a convolutional layer or a fully connected layer.
[0151] In a possible implementation manner, the first processing module 11 is further configured to receive a trajectory prediction model sent by a server and deploy the trajectory prediction model.
[0152] The trajectory prediction device for traffic participants provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and the implementation principles and beneficial effects are similar, which will not be elaborated here.
[0153] Figure 8 It is a schematic structural diagram of a model training device provided by an embodiment of the present application. Please refer to Figure 8 The model training device 20 may include:
[0154] An acquisition module 21, configured to acquire a plurality of training samples, each training sample including a perception vector and a driving trajectory corresponding to the perception vector, where the perception vector is used to represent the characteristics of traffic participants and the traffic environment;
[0155] A processing module 22, configured to train a deep learning model according to the plurality of training samples to obtain a trajectory prediction model.
[0156] In a possible implementation manner, the processing module 22 is further configured to send the trajectory prediction model to at least one vehicle.
[0157] The model training device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and the implementation principles and beneficial effects are similar, which will not be elaborated here.
[0158] Figure 9 It is a schematic structural diagram of an electronic device provided by the present application. Please refer to Figure 9, the electronic device 30 provided in the embodiments of the present application includes: at least one processor 31 and a memory 32. Optionally, the device 30 further includes a communication component 33. Among them, the processor 31, the memory 32, and the communication component 33 are connected through a bus 34.
[0159] In a specific implementation process, at least one processor 31 executes the computer-executable instructions stored in the memory 32, so that at least one processor 31 executes the above-mentioned method.
[0160] The electronic device provided in the embodiments of the present application can be the vehicle or the server in the above embodiments. When the electronic device is a vehicle, it can be used to execute the method provided in the method embodiment related to the vehicle side; when the electronic device is a server, it can be used to execute the method provided in the method embodiment related to the server side; the specific implementation process of the processor 31 can refer to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0161] In the above embodiments, it should be understood that the processor may be a central processing unit (Central Processing Unit, CPU for short), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP for short), application specific integrated circuits (Application Specific Integrated Circuit, ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by a hardware processor, or executed and completed by a combination of hardware and software modules in the processor.
[0162] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0163] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0164] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0165] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the above method is implemented.
[0166] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0167] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.
[0168] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0169] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0170] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0171] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., various media that can store program codes.
[0172] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROMs, RAMs, magnetic disks, or optical discs, etc., various media that can store program codes.
[0173] Finally, it should be noted that: after considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptive changes of the present invention. These variations, uses, or adaptive changes follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A method for predicting the trajectory of a traffic participant, characterized in that, Applied to a vehicle, the method includes: Obtain perception data, where the perception data includes image data and / or point cloud data of traffic participants; Perform inference on the perception data through the backbone network of the perception model to obtain a perception vector; the perception vector is used to represent the characteristics of the traffic participants and the traffic environment; Perform predictive inference on the perception vector to obtain the predicted trajectory of the traffic participant.
2. The method according to claim 1, characterized in that, The performing predictive inference on the perception vector to obtain the predicted trajectory of the traffic participant includes: According to the trajectory prediction model, perform predictive inference on the perception vector to obtain the predicted trajectory of the traffic participant; Wherein, the trajectory prediction model is a model for performing trajectory prediction based on a perception vector obtained by training a deep learning model.
3. The method according to claim 1, characterized in that, The performing inference on the perception data through the backbone network of the perception model to obtain a perception vector includes: Determine a target layer in the backbone network of the perception model, where the target layer is the feature processing layer in the backbone network that outputs the richest semantic features; Perform forward propagation calculation on the perception data to obtain an intermediate vector corresponding to the perception data; Perform inference on the intermediate vector through the target layer to obtain the perception vector.
4. The method according to claim 3, wherein The feature processing layer is a convolutional layer or a fully connected layer.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Receive the trajectory prediction model sent by the server and deploy the trajectory prediction model.
6. A model training method, characterized in that, Applied to a server, the method includes: Obtain a plurality of training samples, where each training sample includes a perception vector and the driving trajectory corresponding to the perception vector, and the perception vector is used to represent the characteristics of the traffic participants and the traffic environment; Train a deep learning model according to the plurality of training samples to obtain a trajectory prediction model.
7. The method according to claim 6, wherein The method further includes: Send the trajectory prediction model to at least one vehicle.
8. A trajectory prediction device for a traffic participant, characterized in that, The device includes: A first processing module, configured to obtain perception data, where the perception data includes image data and / or point cloud data of traffic participants; A second processing module, configured to perform inference on the perception data through the backbone network of the perception model to obtain a perception vector; the perception vector is used to represent the characteristics of the traffic participants and the traffic environment; A prediction module, configured to perform predictive inference on the perception vector to obtain the predicted trajectory of the traffic participant.
9. A model training device, characterized in that, The device includes: An acquisition module, configured to obtain a plurality of training samples, where each training sample includes a perception vector and the driving trajectory corresponding to the perception vector, and the perception vector is used to represent the characteristics of the traffic participants and the traffic environment; A processing module, configured to train a deep learning model according to the plurality of training samples to obtain a trajectory prediction model.
10. A vehicle, characterized in that, Includes: A processor, a memory, and a communication interface; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1-5.
11. A server, characterized in that, Includes: A processor, a memory, and a communication interface; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to claim 6 or 7.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, are used to implement the method according to any one of claims 1-7.