Training method and device of trajectory prediction model, equipment, medium and program product

By extracting image samples from driving videos and processing them through affine transformation, and combining with truth-value training of trajectory prediction models, the problem of poor generalization ability of existing models in complex scenarios is solved, and a more efficient and safe trajectory prediction is achieved.

CN120182745APending Publication Date: 2025-06-20NAVINFO SMART DRIVING (BEIJING) TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510221530.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The trajectory prediction model in existing autonomous driving systems performs poorly in complex road scenarios and vehicle offset scenarios, and lacks generalization capabilities and robustness, resulting in safety and accuracy of predicted trajectories.

Method used

By obtaining the driving video from the vehicle perspective, extracting the image as the first sample, and generating the second sample through affine transformation, determining the sample truth value with the driving video, and training the trajectory prediction model. This method not only improves the training efficiency of the model, but also enhances the model's ability to handle vehicle offset scenarios.

Benefits of technology

The generalization ability and robustness of the trajectory prediction model are improved, so that it can provide more accurate and safe prediction trajectory in complex road environments and vehicle offset scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182745A_ABST
    Figure CN120182745A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a trajectory prediction model training method and device, equipment, a medium and a program product. The method comprises the following steps: acquiring a driving video of a vehicle view angle, extracting an image from the video as a first sample, performing affine transformation processing on the first sample to obtain a second sample, determining a sample truth value in combination with the driving video, and performing model training by using the first sample and the second sample to obtain a trajectory prediction model. A large number of continuous image frames can be extracted from the driving video, samples for training can be quickly obtained, truth values can be determined according to subsequent image frames of the samples, and the model training efficiency is improved. Besides, the second sample obtained by affine transformation of the first sample can simulate the vehicle offset condition, and when the second sample is used in the model training process, the vehicle offset scene processing capability of the model can be improved, and the generalization capability and robustness of the model can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving, and in particular to a training method, device, equipment, medium and program product for a trajectory prediction model. Background Art

[0002] With the development of artificial intelligence, autonomous driving technology has become more and more mature and has become a hot topic in the automotive field. With the environmental information collected by sensors and other tools, the autonomous driving system can control the vehicle to drive along the planned trajectory, thereby replacing the operation of human drivers to a certain extent.

[0003] At present, autonomous driving technology mostly uses training models to predict and plan the future trajectory of the vehicle. The main principle is to enable the model to learn the driving trajectory of human drivers through supervised learning and imitation learning. However, this model has different performances in actual applications, lacks generalization ability and robustness, and has problems such as insufficient safety of the driving trajectory predicted by the model when it is out of specific usage scenarios. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus, device, medium and program product for training a trajectory prediction model, so as to achieve the effect of improving the generalization ability and robustness of the trajectory prediction model.

[0005] In a first aspect, an embodiment of the present application provides a method for training a trajectory prediction model, comprising:

[0006] Obtain driving video from the vehicle's perspective;

[0007] Extracting an image from the driving video as a first sample;

[0008] Performing an affine transformation on the first sample, and using the affine transformation result of the first sample as the second sample;

[0009] Determining true values ​​of the first sample and the second sample according to the driving video;

[0010] A trajectory prediction model is trained based on the first sample, the second sample, and the true value.

[0011] In a possible implementation manner, determining the true values ​​of the first sample and the second sample according to the driving video includes:

[0012] Taking the N frames of images after the first sample in the driving video as true value images, and extracting the driving trajectory according to the true value images as the first true value corresponding to the first sample, where N is a positive integer;

[0013] Use the affine transformation result of the true value image as the second true value corresponding to the second sample, or use the result of transforming the driving trajectory according to the parameters of the affine transformation as the second true value corresponding to the second sample.

[0014] In a possible implementation manner, training a trajectory prediction model based on the first sample, the second sample, and the true value includes:

[0015] Input the first sample into an initial model to be trained, obtain a first predicted trajectory output by the initial model, and use the difference between the first predicted trajectory and the true value of the first sample as the first training loss;

[0016] Input the second sample into the initial model to be trained, obtain a second predicted trajectory output by the initial model, and use the difference between the second predicted trajectory and the true value of the second sample as the second training loss;

[0017] Adjust the parameters of the initial model to reduce the first training loss and the second training loss. After the adjustment is completed, obtain a trained trajectory prediction model.

[0018] In a possible implementation manner, training a trajectory prediction model based on the first sample, the second sample, and the true value further includes:

[0019] Obtain the feature distribution of the last layer when the trajectory prediction model takes the training sample as the input, denoted as the training distribution, where the training sample is the first sample or the second sample;

[0020] Denote the feature distribution of the true value corresponding to the training sample as the target distribution;

[0021] Calculate the divergence value of the training distribution relative to the target distribution through the asymmetric divergence algorithm, and use the divergence value as the third training loss to adjust the parameters of the trajectory prediction model.

[0022] In a possible implementation manner, training a trajectory prediction model based on the first sample, the second sample, and the true value further includes:

[0023] Select target points from the second predicted trajectory, where the second predicted trajectory represents the predicted position of the vehicle within the next t seconds, and the target point represents the predicted position of the vehicle at the r-th second, where t and r are positive numbers, and t > r;

[0024] Determine the actual position of the vehicle at the r-th second according to the true value of the second sample;

[0025] Adjust the parameters of the trajectory prediction model to reduce the position difference between the predicted position and the actual position.

[0026] In a possible implementation, extracting an image from the driving video as a first sample includes:

[0027] Obtaining the camera parameters during the acquisition process of the driving video;

[0028] Selecting a target image from the driving video;

[0029] Converting the target image into a preset coordinate system according to the camera parameters, and obtaining the first sample after conversion.

[0030] In a possible implementation, it further includes:

[0031] Adding random Gaussian noise during the process of training the trajectory prediction model.

[0032] In a possible implementation, the affine transformation includes at least one of rotation, scaling, translation, and shearing.

[0033] In a second aspect, an embodiment of the present application provides a training device for a trajectory prediction model, including:

[0034] An acquisition module, configured to acquire a driving video from the vehicle perspective;

[0035] A first extraction module, configured to extract an image from the driving video as a first sample;

[0036] A second extraction module, configured to perform an affine transformation on the first sample, and use the affine transformation result of the first sample as a second sample;

[0037] A ground truth processing module, configured to determine the ground truth of the first sample and the second sample according to the driving video;

[0038] A model training module, configured to train a trajectory prediction model based on the first sample, the second sample, and the ground truth.

[0039] In a possible implementation, the ground truth processing module is further configured to: use the last N frames of images of the first sample in the driving video as ground truth images, and extract the driving trajectory according to the ground truth images as the first ground truth corresponding to the first sample, where N is a positive integer; use the affine transformation result of the ground truth image as the second ground truth corresponding to the second sample, or use the result of transforming the driving trajectory according to the parameters of the affine transformation as the second ground truth corresponding to the second sample.

[0040] In a possible implementation, the model training module is further configured to: input the first sample into an initial model to be trained, obtain a first predicted trajectory output by the initial model, and use the difference between the first predicted trajectory and the ground truth of the first sample as a first training loss; input the second sample into the initial model to be trained, obtain a second predicted trajectory output by the initial model, and use the difference between the second predicted trajectory and the ground truth of the second sample as a second training loss; adjust the parameters of the initial model to reduce the first training loss and the second training loss, and after the adjustment is completed, obtain a trained trajectory prediction model.

[0041] In a possible implementation, the model training module is further configured to: obtain the feature distribution of the last layer when the trajectory prediction model takes a training sample as input, denoted as a training distribution, where the training sample is the first sample or the second sample; denote the feature distribution of the ground truth corresponding to the training sample as a target distribution; calculate the divergence value of the training distribution relative to the target distribution through an asymmetric divergence algorithm, and use the divergence value as a third training loss to adjust the parameters of the trajectory prediction model.

[0042] In a possible implementation, the model training module is further configured to: select a target point from the second predicted trajectory, where the second predicted trajectory represents the predicted position of the vehicle within the next t seconds, and the target point represents the predicted position of the vehicle at the r-th second, t and r are positive numbers, and t > r; determine the actual position of the vehicle at the r-th second according to the ground truth of the second sample; adjust the parameters of the trajectory prediction model to reduce the position difference between the predicted position and the actual position.

[0043] In a possible implementation, the first extraction module is further configured to: obtain the camera parameters during the acquisition process of the driving video; select a target image from the driving video; convert the target image into a preset coordinate system according to the camera parameters, and obtain a first sample after the conversion.

[0044] In a possible implementation, the model training module is further configured to: add random Gaussian noise during the process of training the trajectory prediction model.

[0045] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0046] The memory stores computer execution instructions;

[0047] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0048] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above first aspect and / or various possible implementation manners of the first aspect when executed by a processor.

[0049] Fifthly, an embodiment of the present application provides a computer program product, including a computer program, which implements the above first aspect and / or various possible implementation manners of the first aspect when executed by a processor.

[0050] The training method, device, equipment, medium and program product of the trajectory prediction model provided by the embodiments of the present application can extract images from the driving video as the first samples by acquiring the driving video from the vehicle perspective. The first samples can be processed by affine transformation to obtain the second samples. Then, after determining the sample truth values in combination with the driving video, the first samples and the second samples can be used for model training to obtain the trajectory prediction model. A large number of continuous image frames can be extracted from the driving video, which can not only quickly obtain the samples for training, but also determine the truth values according to the subsequent image frames of the samples, improving the model training efficiency. In addition, the second samples obtained by affine transformation of the first samples can simulate the situation of vehicle offset. Using them in the model training process can improve the model's ability to process vehicle offset scenarios and enhance the generalization ability and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0052] Figure 1 It is a schematic diagram of an application scenario provided by the present application;

[0053] Figure 2 It is a flowchart of a training method of a trajectory prediction model provided by the present application Figure 1 ;

[0054] Figure 3 It is a flowchart of a training method of a trajectory prediction model provided by the present application Figure 2 ;

[0055] Figure 4 It is a flowchart of a training method of a trajectory prediction model provided by the present application Figure 3 ;

[0056] Figure 5 It is a flowchart of a training method of a trajectory prediction model provided by the present application Figure 4 ;

[0057] Figure 6Structural schematic diagram of a training device for a trajectory prediction model provided by this application;

[0058] Figure 7 Structural schematic diagram of an electronic device provided by this application.

[0059] Through the above-mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Specific embodiments

[0060] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0061] With the development of artificial intelligence technology, especially in the field of autonomous driving, the demand for end-to-end models capable of handling complex real-world environments is increasing. An autonomous driving system can perceive the vehicle's external environment and use a model to predict the next driving trajectory, and then execute corresponding driving operations according to the driving trajectory.

[0062] Currently, the training of end-to-end models in the field of autonomous driving mainly relies on supervised learning and imitation learning. These methods usually require a large amount of labeled data, such as video streams and corresponding driving trajectories, to train the model to recognize road information and plan trajectories.

[0063] Figure 1 Schematic diagram of an application scenario provided by this application. As Figure 1 shown, this application can be specifically used to train a model required for an autonomous driving system to perform trajectory prediction. For a vehicle equipped with an autonomous driving system, the vehicle can collect data on the surrounding environment of the vehicle through its own cameras and other sensors. The autonomous driving system can process these data and use the model to predict and plan the next driving trajectory of the vehicle, and the control module can execute operations such as lane changes and turns to make the vehicle drive according to the trajectory output by the model.

[0064] However, the trajectory prediction model (i.e., the model used by the autonomous driving system to predict the future driving trajectory of a vehicle) has varying performances in actual road tests. It performs acceptably in some relatively conventional road scenarios. However, the real environment is complex and changeable. In complex road scenarios, especially those involving unconventional situations such as vehicle deviation, there are significant deficiencies in the accuracy and safety of the trajectories predicted by such models. They are unable to handle diverse scenarios and have poor generalization ability.

[0065] Combined with the above scenarios, it can be seen that the trajectory prediction model used in autonomous driving currently still has technical problems such as poor generalization ability and insufficient robustness.

[0066] The inventors have found through research that the reason for the above technical problems in the trajectory prediction model is that the training of this model relies on real driving videos. Usually, the videos containing road scenes collected by the vehicle's own camera during the actual driving of the driver are used as training data, resulting in the performance of the trained model being restricted by the training data. When the vehicle faces some road environments that are significantly different from the training data, the model performance drops significantly, and it is difficult to provide high-quality predicted trajectories. For example, the driving videos used for model training usually record the normal driving process of the driver, and the scenarios of vehicle deviation are missing in the training data. The model cannot learn environmental information such as the lane line positions in such scenarios, nor can it learn the correct trajectory that should be returned after deviation, resulting in the trained model lacking the ability to handle vehicle deviation. The model has poor performance in dealing with scenarios where the vehicle environment changes or the result of a certain frame in the middle is offset. It is possible that after the vehicle deviates due to various noises, the subsequent results become more and more deviated, even leading to accidents.

[0067] Based on the above findings, the inventors propose a new training method. Select vehicle driving videos as training data and add samples of vehicle deviation and the return to the correct driving position after deviation to the training data, so that the model can learn the ability to handle deviation scenarios and improve the generalization ability of the model.

[0068] Specifically, methods such as using a simulator or a 3D rendering engine can be used to simulate vehicle deviation scenarios in real driving to obtain corresponding deviation samples. Affine transformation can also be performed on the original training data to generate new samples. These new samples will be offset compared to the original perspective, so they can also be used as deviation samples. Compared with the deviation samples obtained by simulating using a simulator or a 3D rendering engine, the samples obtained by performing affine transformation on the original training data are more in line with the data distribution of the real world. Therefore, samples representing vehicle deviation can be obtained by performing affine transformation on the original training data, and the ground truth of the samples can be determined based on the original training data, thereby improving the model's ability to handle vehicle deviation scenarios and enhancing the generalization ability and robustness of the model.

[0069] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0070] Figure 2 Schematic flow of a method for training a trajectory prediction model provided by this application Figure 1 , as Figure 2 shown, the method includes:

[0071] Step S201, obtain a driving video from the vehicle's perspective.

[0072] Among them, the driving video can be a video collected by vehicle sensors with the vehicle as the first perspective.

[0073] In some possible implementation manners, driving videos under different shooting conditions can be collected. The shooting conditions can refer to factors related to the vehicle, and these factors may involve changes in shooting height. The shooting conditions can also refer to factors related to the vehicle's surrounding environment (such as different weather conditions, different time periods, different traffic conditions, etc.). For example, the fields of view of a sedan and a large truck are different. To take into account different vehicle types, driving videos taken at different heights by a camera at a preset angle can be collected. By collecting driving videos under different shooting conditions, the diversity of samples can be increased, and the generalization ability of the model can be improved.

[0074] Step S202, extract images from the driving video as the first samples.

[0075] In the embodiments of this application, the driving video can include an environmental picture observed from the vehicle's perspective as the first perspective. Images frames related to the vehicle's driving can be selected from the video as the first samples for training the model. For example, in the driving video, the image frames corresponding to the vehicle parked for a long time have little effect on predicting the driving trajectory, while the image frames corresponding to the vehicle starting and driving have a greater effect on predicting the trajectory. The first samples can be selected from the image frames related to the vehicle starting and driving.

[0076] In some possible implementation manners, preprocessing such as data cleaning or data augmentation can be performed on the images used as the first samples to enhance the ability of the model. For example, operations such as scaling and cropping can be performed on the images, and different lighting conditions and weather conditions can also be simulated by processing the images. Through these methods, the generalization ability of the model can be improved.

[0077] Step S203, perform an affine transformation on the first samples, and use the affine transformation result of the first samples as the second samples.

[0078] Among them, the affine transformation includes at least one of rotation, scaling, translation, and shearing.

[0079] Specifically, an affine transformation operation can be performed on the image as the first sample to perform a coordinate system conversion on the elements in the image. During the conversion process, the relative positions and attributes of the coordinate points do not change, and the converted image can be used as the second sample. For example, the first sample may include elements such as lane lines and obstacles. By performing a linear transformation and translation on the first sample, the second sample can be obtained. The lane lines in the second sample can still maintain a parallel relationship, but the elements in the second sample have changed relative to the original perspective, thereby simulating the situation where the vehicle deviates from the lane.

[0080] In this way, a second sample that can simulate the vehicle deviation scenario can be obtained. Using the second sample for the training of the trajectory prediction model can not only improve the adaptability of the model to different perspectives and positions, but also enhance the generalization ability of the model when facing complex road environments such as vehicle deviation. It should be noted that there can be several first samples, and it is not necessary to perform an affine transformation on all the first samples. The numbers of the first sample and the second sample can be custom-set according to the needs of model training (for example, the ratio of the number of the first sample to the number of the second sample can be 10:1), and the present application does not limit the specific numbers of the first sample and the second sample.

[0081] Step S204, determine the true values of the first sample and the second sample according to the driving video.

[0082] Among them, the true value refers to the ideal output value used for training and evaluating the model, which can be the true label annotated for the sample. The true value is the benchmark for the model to learn, and it can be used to guide the model to adjust its parameters to improve the prediction accuracy.

[0083] In the embodiment of the present application, the true value of the trajectory prediction model can be the correct driving trajectory.

[0084] Exemplarily, if the first sample is an image frame in the driving video, the actual driving trajectory of the vehicle can be extracted based on this frame image and the following several frame images as the true value of the first sample, or the actual driving trajectory can be extracted in advance based on multiple consecutive frame images. If there is an image selected as the first sample among the multiple frame images, the following trajectory corresponding to the first sample can be intercepted as the true value of the first sample.

[0085] Exemplarily, if the second sample is obtained by performing an affine transformation on the first sample, the true value of the first sample can be processed according to the same affine transformation to obtain the true value of the second sample.

[0086] Step S205, train the trajectory prediction model based on the first sample, the second sample, and the true value.

[0087] Specifically, the first sample and the second sample can be used as the training dataset, and a neural network based on the attention mechanism can be selected as the initial model. The samples in the dataset are input into the initial model, and the training objective is to reduce the difference between the predicted value output by the model and the true value of the sample. Through continuous iteration by adjusting the model parameters, until the difference between the predicted value output by the model and the true value of the sample meets the preset standard, the iteration can be stopped, and the model under the corresponding parameters is used as the trained trajectory prediction model.

[0088] In the above embodiment, by acquiring the driving video from the vehicle perspective, images can be extracted from the video as the first sample. After affine transformation processing of the first sample, the second sample can be obtained. Then, after determining the true value of the sample in combination with the driving video, the first sample and the second sample can be used for model training to obtain the trajectory prediction model. A large number of continuous image frames can be extracted from the driving video, which can not only quickly obtain the samples for training, but also determine the true value according to the subsequent image frames of the samples, improving the model training efficiency. In addition, the second sample obtained by affine transformation of the first sample can simulate the situation of vehicle offset. Using it in the model training process can improve the model's ability to handle vehicle offset scenarios and enhance the generalization ability and robustness of the model.

[0089] In one embodiment, extracting images from the driving video as the first sample includes:

[0090] Obtaining the camera parameters during the acquisition process of the driving video; selecting the target image from the driving video; and converting the target image to a preset coordinate system according to the camera parameters. After conversion, the first sample is obtained.

[0091] Among them, the camera parameters may include the internal and external parameters of the shooting device used to shoot the driving video.

[0092] Exemplarily, the target images related to vehicle startup and driving can be extracted from the driving video, and the extracted images are converted to the same coordinate system required for model training. At the same time, according to the coordinate transformation relationship, the external parameters of the camera center position and the rear axle center position of different vehicles are calculated to obtain a virtual camera coordinate system position, and the images are converted to the virtual camera coordinate system.

[0093] In one embodiment, determining the true values of the first sample and the second sample according to the driving video includes:

[0094] First, taking the last N frames of images of the first sample in the driving video as the true value images, and extracting the driving trajectory according to the true value images as the first true value corresponding to the first sample. Where N is a positive integer.

[0095] Specifically, the actual driving trajectory of the vehicle can be extracted from consecutive image frames in the driving video, including information such as the position, speed, and direction of the vehicle. If an image frame is used as the first sample, then the N frames of images after this image frame can be determined as the ground-truth images, and the trajectory part corresponding to the ground-truth images is used as the ground-truth of the first sample.

[0096] Second, the affine transformation result of the ground-truth image is used as the second ground-truth corresponding to the second sample, or the result of transforming the driving trajectory according to the parameters of the affine transformation is used as the second ground-truth corresponding to the second sample.

[0097] Exemplarily, for any first sample, after performing an affine transformation on the first sample, the operation parameters of the affine transformation can be recorded, and the N ground-truth images of the first sample in the driving video can also be subjected to an affine transformation operation according to the operation parameters to obtain N affine transformation results, and the vehicle driving trajectory is extracted based on these N affine transformation results as the ground-truth of the second sample.

[0098] In some possible implementation manners, the corresponding transformation matrix can also be determined according to the process of performing an affine transformation on the first sample, and the coordinate points in the subsequent actual driving trajectory of the image frame of the first sample are calculated according to the transformation matrix, and the obtained transformed coordinate points form the corresponding driving trajectory as the ground-truth of the second sample.

[0099] In the above embodiments, by processing the subsequent image frames of the first sample in the driving video, the actual driving trajectory of the vehicle can be extracted as the ground-truth of the first sample, and the ground-truth of the second sample can be obtained according to the parameters related to the affine transformation. The acquisition of the ground-truth can be automatically implemented without manual annotation of the sample ground-truth, so that the training process of the trajectory prediction model can be freed from manual annotation and the model training efficiency can be improved.

[0100] In one embodiment, as Figure 3 shown, training a trajectory prediction model based on the first sample, the second sample, and the ground-truth includes:

[0101] Step S301: Input the first sample into the initial model to be trained, obtain the first predicted trajectory output by the initial model, and use the difference between the first predicted trajectory and the ground-truth of the first sample as the first training loss.

[0102] Step S302: Input the second sample into the initial model to be trained, obtain the second predicted trajectory output by the initial model, and use the difference between the second predicted trajectory and the ground-truth of the second sample as the second training loss.

[0103] Step S303: Adjust the parameters of the initial model to reduce the first training loss and the second training loss. After the adjustment is completed, the trained trajectory prediction model is obtained.

[0104] In the embodiments of the present application, a deep learning algorithm can be used to train the trajectory prediction model. For example, a convolutional neural network (CNN) and a self-attention mechanism can be applied to encode and decode the image features corresponding to the first sample and the second sample. The convolutional neural network can extract spatial features from images, while the self-attention mechanism can capture long-range dependencies in sequence data. Through deep learning, the model can learn how to extract key feature information from the input images and thus obtain the predicted trajectory. This training method enables the model to adapt to different road environments and traffic conditions and improves the accuracy of trajectory prediction.

[0105] In addition, vehicle driving has temporality, and the self-attention module can efficiently extract the temporal features during vehicle driving. Optionally, by selecting appropriate consecutive frame data from the driving video, it can be spliced onto the channel, enabling the model to give the processing results of multiple future temporal frames in one inference. This method can not only improve the inference efficiency of the model but also make the results combine the image transformations of the front and rear frames to better give the predicted trajectory. Through temporal training, the model can learn the motion trends and trajectory changes of the vehicle, improving the continuity and stability of trajectory prediction.

[0106] In one embodiment, as Figure 4 shown, training the trajectory prediction model based on the first sample, the second sample, and the ground truth further includes:

[0107] Step S401: Obtain the feature distribution of the last layer of the trajectory prediction model when the training sample is used as the input, denoted as the training distribution.

[0108] Wherein, the training sample is the first sample or the second sample.

[0109] The trajectory prediction model adopts a neural network structure. The model can include multiple intermediate layers and a final output layer. Among them, the intermediate layers can be used to process the features extracted by the model. The intermediate layers finally input the processed features into the last layer (i.e., the output layer) of the model, and the output layer can be used to output the predicted trajectory.

[0110] Step S402: Denote the feature distribution of the ground truth corresponding to the training sample as the target distribution.

[0111] Exemplarily, the training sample can be an image frame selected from a driving video. Based on this image frame, the true driving trajectory of the vehicle in several subsequent frames of the driving video can be extracted as the ground truth of the training sample. When using the training sample as the input of the trajectory prediction model, the input value of the last layer of the trajectory prediction model is the feature vector obtained by processing the training sample through multiple intermediate layers. The distribution of these feature vectors can be expressed in the form of a probability distribution, denoted as the training distribution. Referring to the distribution of the feature vectors of the last layer of the trajectory prediction model, the ground truth can also be expressed as features in the corresponding form, and the ground truth feature distribution is denoted as the target distribution.

[0112] Step S403: Calculate the divergence value of the training distribution relative to the target distribution through the asymmetric divergence algorithm, and use the divergence value as the third training loss to adjust the parameters of the trajectory prediction model.

[0113] Among them, the asymmetric divergence algorithm can be the KL divergence algorithm. KL (Kullback-Leibler Divergence) divergence is an asymmetric measure that can be used to measure the difference between two distributions.

[0114] In some possible implementation manners, the KL divergence can be used to construct a loss function, and the divergence value is calculated to measure the difference between the feature distribution (i.e., the above-mentioned training distribution) processed by the trajectory prediction model and the feature distribution of the ground truth. When the divergence value is large, it indicates that the difference degree between the two is large; when the divergence value is small, it indicates that the difference between the two is small.

[0115] During the training process of the trajectory prediction model, on the basis of the first training loss and the second training loss, the divergence value can be added as the third training loss, and the model parameters are adjusted through backpropagation to reduce the first training loss, the second training loss, and the third training loss.

[0116] The training distribution is the basis for the last layer of the trajectory prediction model to make a trajectory prediction. Taking this divergence value as the model training loss, the model parameters are adjusted to reduce this divergence value. In this way, the feature distribution of the actual trajectory can be used as the target for approximating the feature distribution of the last layer of the trajectory prediction model. The smaller the divergence value, the lower the information loss degree when using the model feature distribution to approximate the actual distribution. The training of the trajectory prediction model is based on the learning of the actual trajectory of the vehicle in the driving video. By using the asymmetric divergence value as the training loss, the feature distribution of the last layer of the trajectory prediction model can better meet the training requirements, making the model output feature distribution closer to the Gaussian distribution and improving the prediction accuracy of the model.

[0117] In one embodiment, as Figure 5 shown, when training the trajectory prediction model based on the first sample, the second sample, and the ground truth, it further includes:

[0118] Step S501: Select a target point from the second predicted trajectory.

[0119] Wherein, the second predicted trajectory represents the predicted position of the vehicle within the next t seconds, the target point represents the predicted position of the vehicle at the r-th second, t and r are positive numbers, and t > r.

[0120] Step S502: Determine the actual position of the vehicle at the r-th second according to the ground truth of the second sample.

[0121] Step S503: Adjust the parameters of the trajectory prediction model to reduce the position difference between the predicted position and the actual position.

[0122] Exemplarily, the coordinates of the vehicle at the r-th second can be determined according to the second predicted trajectory, denoted as the predicted position; the coordinates of the vehicle at the r-th second can be determined according to the ground truth of the second sample, denoted as the actual position; calculate the coordinate difference between the actual position and the predicted position, and reduce this difference by adjusting the parameters of the trajectory prediction model.

[0123] The predicted trajectory output by the trajectory prediction model represents the prediction result of the model on the position of the vehicle within the next t seconds, and the ground truth, i.e., the true trajectory, represents the actual position of the vehicle within the next t seconds during the actual driving process represented by the driving video. Among them, the t seconds are based on the moment when the image frame corresponding to the second sample is located in the driving video.

[0124] In the embodiments of the present application, the second sample is used to simulate the situation of vehicle deviation. If the vehicle deviates, it is not only necessary to predict the next ideal trajectory, but also to make the vehicle return to the correct path as quickly as possible. If only the overall difference between the second predicted trajectory and the second ground truth is calculated as the training loss, although the predicted trajectory may generally match the actual trajectory in the driving video, there may be a situation where it cannot return to the correct path in a short time. For example, after the vehicle deviates, the model gives a predicted trajectory within the next t seconds (such as within 10 seconds), and the vehicle needs to return to the center line of the lane at the r-th second (such as the 5th second) according to this predicted trajectory, and still deviates from the center line before the 5th second, which poses a safety hazard. Therefore, in the above embodiments, by reasonably setting the value of r, the position accuracy of the predicted trajectory at the r-th second can be used as the training loss of the model, so that the model can learn to return to the correct path in a short time, enhance the ability of the model to correct vehicle deviation, and improve the safety of autonomous driving.

[0125] In one embodiment, it further includes: adding random Gaussian noise during the process of training the trajectory prediction model.

[0126] In some possible implementation manners, in order to avoid the feature distortion of the image road information caused by the model overlearning the affine transformation, during the training process of the trajectory prediction model, random Gaussian noise calculation can be introduced. This method enables the model to correct the offset only when facing feature distortion, and compresses the content that can be represented by the last-layer feature vector by adding random Gaussian noise.

[0127] Exemplarily, the information capacity of the Gaussian channel can be calculated using the Shannon formula, and the signal-to-noise ratio can be dynamically adjusted to ensure that the model does not learn the feature distortion caused by the affine transformation, thereby enhancing the robustness of the model to noise and distortion and making it more stable and reliable in practical applications.

[0128] In some possible implementation manners, during the training process of the trajectory prediction model, data from different sensors, such as radar, lidar (LiDAR), camera, etc., can also be combined to enhance the model's perception ability of the environment.

[0129] In some possible implementation manners, the generalization ability of the model can be increased through generative adversarial networks (GANs), enabling the model to handle a wider range of deviation situations.

[0130] For example, the generator network and the discriminator network can be designed through the GANs (Generative Adversarial Networks) algorithm. The first sample and the second sample are input into the generator network after adding offset noise, so that the generator network outputs a simulated image simulating the offset scenario, and the simulated image is input into the discriminator network, so that the discriminator network outputs the result of whether the simulated image conforms to the real offset scenario. Among them, adding offset noise can refer to randomly changing the lateral offset distance, heading angle, etc. on the basis of the first sample or the second sample.

[0131] Through adversarial training, the generator network can generate samples more conforming to the actual offset scenario. Then, the samples generated by the generator network are combined with the second sample obtained through affine transformation, which can enhance the original data set and make it cover the scenarios that are difficult to simulate by affine transformation. Using the enhanced data set to train the vehicle trajectory prediction model can increase the generalization ability of the model and better handle the vehicle offset situation.

[0132] In some possible implementation manners, the second sample can also be processed through a depth estimation algorithm to obtain a more realistic transformed image, reduce the feature distortion caused by the affine transformation, so that the model does not focus on learning the distortion but learns more road image information.

[0133] For the second sample obtained by affine transformation of the first sample, the depth information of the image of the first sample can be obtained using a depth estimation algorithm, and then the second sample can be geometrically corrected and detail-enhanced according to this depth information to make it closer to the real scene represented by the first sample. For example, the depth information can be used to geometrically correct the second sample to adjust the perspective distortion caused by the affine transformation, making the image more conform to the geometric structure of the real world; or, the second sample can be detail-enhanced according to the depth information, such as appropriately adjusting the details of distant and nearby objects to improve the realism of the image.

[0134] In one embodiment, after the training of the trajectory prediction model reaches certain requirements, the model can also be evaluated in a simulation environment. This can evaluate the performance of the model without actually driving a vehicle. According to the evaluation results, we can adjust the model parameters and training strategies to optimize the model performance. This method can not only improve the efficiency of evaluation but also reduce the risks and costs of testing.

[0135] In one embodiment, after the training of the trajectory prediction model reaches certain requirements, the model can also be deployed to actual roads for testing. In the actual road test, the performance data of the model in the real environment can be collected, and these data include the adaptability of the model to different road conditions, traffic conditions, and weather conditions. Through the actual road test, the generalization ability and robustness of the model in actual applications can be evaluated, and the model can be optimized accordingly.

[0136] According to the feedback from the actual road test, the model performance can be continuously iteratively optimized. This iterative optimization can include adjusting the model structure, optimizing the loss function, improving the training strategy, etc. Through continuous iterative optimization, the generalization ability and robustness of the model can be improved, enabling it to operate stably in various complex road environments, thereby improving the safety and reliability of the autonomous driving system.

[0137] In one embodiment, the present application also provides a trajectory prediction method applicable to autonomous driving. This trajectory prediction method can deploy the trajectory prediction model trained based on the above embodiments to the autonomous driving system of the vehicle. The autonomous driving system can collect the surrounding images of the vehicle, perceive the external environment of the vehicle, and can transmit data such as images representing the external environment to the trajectory prediction model to obtain the predicted trajectory output by the trajectory prediction model, and control the vehicle to drive according to this predicted trajectory.

[0138] Figure 6 The structural schematic diagram of a training device for a trajectory prediction model provided by the present application is as Figure 6 shown. The training device 600 for the trajectory prediction model provided in this embodiment includes:

[0139] An acquisition module 601, configured to acquire a driving video from the perspective of a vehicle;

[0140] A first extraction module 602, configured to extract an image from the driving video as a first sample;

[0141] A second extraction module 603, configured to perform an affine transformation on the first sample, and use the affine transformation result of the first sample as a second sample;

[0142] A ground truth processing module 604, configured to determine the ground truth of the first sample and the second sample according to the driving video;

[0143] A model training module 605, configured to train a trajectory prediction model based on the first sample, the second sample, and the ground truth.

[0144] In a possible implementation manner, the ground truth processing module 604 is further configured to: use the subsequent N frames of images of the first sample in the driving video as ground truth images, and extract a driving trajectory according to the ground truth images as the first ground truth corresponding to the first sample, where N is a positive integer; use the affine transformation result of the ground truth images as the second ground truth corresponding to the second sample, or use the result of transforming the driving trajectory according to the parameters of the affine transformation as the second ground truth corresponding to the second sample.

[0145] In a possible implementation manner, the model training module 605 is further configured to: input the first sample into an initial model to be trained, obtain a first predicted trajectory output by the initial model, and use the difference between the first predicted trajectory and the ground truth of the first sample as a first training loss; input the second sample into the initial model to be trained, obtain a second predicted trajectory output by the initial model, and use the difference between the second predicted trajectory and the ground truth of the second sample as a second training loss; adjust the parameters of the initial model to reduce the first training loss and the second training loss, and after the adjustment is completed, obtain a trained trajectory prediction model.

[0146] In a possible implementation manner, the model training module 605 is further configured to: obtain the feature distribution of the last layer when the trajectory prediction model takes a training sample as an input, denoted as a training distribution, where the training sample is the first sample or the second sample; denote the feature distribution of the ground truth corresponding to the training sample as a target distribution; calculate the divergence value of the training distribution relative to the target distribution through an asymmetric divergence algorithm, and use the divergence value as a third training loss to adjust the parameters of the trajectory prediction model.

[0147] In a possible implementation, the model training module 605 is further configured to: select a target point from the second predicted trajectory, where the second predicted trajectory represents the predicted position of the vehicle within the next t seconds, the target point represents the predicted position of the vehicle at the r-th second, t and r are positive numbers, and t > r; determine the actual position of the vehicle at the r-th second according to the ground truth of the second sample; and adjust the parameters of the trajectory prediction model to reduce the position difference between the predicted position and the actual position.

[0148] In a possible implementation, the first extraction module 602 is further configured to: obtain the camera parameters during the acquisition of the driving video; select a target image from the driving video; and convert the target image into a preset coordinate system according to the camera parameters to obtain a first sample after conversion.

[0149] In a possible implementation, the model training module 605 is further configured to: add random Gaussian noise during the process of training the trajectory prediction model.

[0150] The training device for the trajectory prediction model provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0151] Figure 7 The figure is a schematic structural diagram of an electronic device provided by this application. As Figure 7 shown, the electronic device 70 provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, the device 70 further includes a communication component 703. Among them, the processor 701, the memory 702, and the communication component 703 are connected through a bus 704.

[0152] In a specific implementation process, at least one processor 701 executes the computer execution instructions stored in the memory 702, so that at least one processor 701 executes the above method.

[0153] The specific implementation process of the processor 701 can refer to the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0154] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU for short), or other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or implemented by the combination of the hardware and software modules in the processor.

[0155] The memory may include a high-speed random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk memory.

[0156] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0157] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0158] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0159] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0160] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0161] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings or direct couplings or communication connections shown or discussed between each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0162] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0164] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.

[0165] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical disks that can store program codes.

[0166] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for training a trajectory prediction model, characterized in that: include: Obtain driving video from the vehicle's perspective; Extracting an image from the driving video as a first sample; Performing an affine transformation on the first sample, and using the affine transformation result of the first sample as the second sample; Determining true values ​​of the first sample and the second sample according to the driving video; A trajectory prediction model is trained based on the first sample, the second sample, and the true value.

2. The method according to claim 1, characterized in that The determining the true values ​​of the first sample and the second sample according to the driving video includes: Taking the N frames of images after the first sample in the driving video as true value images, and extracting the driving trajectory according to the true value images as the first true value corresponding to the first sample, where N is a positive integer; An affine transformation result of the true value image is used as the second true value corresponding to the second sample, or a result after the driving trajectory is transformed according to the parameters of the affine transformation is used as the second true value corresponding to the second sample.

3. The method according to claim 1 or 2, characterized in that: The training of the trajectory prediction model based on the first sample, the second sample and the true value includes: Inputting the first sample into an initial model to be trained, obtaining a first predicted trajectory output by the initial model, and taking the difference between the first predicted trajectory and the true value of the first sample as a first training loss; Inputting the second sample into an initial model to be trained to obtain a second predicted trajectory output by the initial model, and taking the difference between the second predicted trajectory and the true value of the second sample as a second training loss; The parameters of the initial model are adjusted to reduce the first training loss and the second training loss. After the adjustment is completed, a trained trajectory prediction model is obtained.

4. The method according to claim 3, characterized in that The trajectory prediction model is trained based on the first sample, the second sample and the true value, further comprising: Obtaining a feature distribution of the last layer of the trajectory prediction model when the training sample is used as input, recorded as training distribution, where the training sample is the first sample or the second sample; Recording the characteristic distribution of the true value corresponding to the training sample as the target distribution; The divergence value of the training distribution relative to the target distribution is calculated by an asymmetric divergence algorithm, and the divergence value is used as a third training loss to adjust the parameters of the trajectory prediction model.

5. The method according to claim 3, characterized in that: The trajectory prediction model is trained based on the first sample, the second sample and the true value, further comprising: Selecting a target point from the second predicted trajectory, the second predicted trajectory represents a predicted position of the vehicle in the next t seconds, and the target point represents a predicted position of the vehicle at the rth second, where t and r are positive numbers and t>r; Determine the actual position of the vehicle at the rth second according to the true value of the second sample; Parameters of the trajectory prediction model are adjusted to reduce the position difference between the predicted position and the actual position.

6. The method according to claim 1 or 2, characterized in that: The step of extracting an image from the driving video as a first sample includes: Obtaining camera parameters of the driving video during the acquisition process; Selecting a target image from the driving video; The target image is converted into a preset coordinate system according to the camera parameters, and a first sample is obtained after the conversion.

7. A training device for a trajectory prediction model, characterized in that: include: An acquisition module, used to acquire driving video from the perspective of the vehicle; A first extraction module, used to extract an image from the driving video as a first sample; A second extraction module, configured to perform an affine transformation on the first sample, and use the affine transformation result of the first sample as the second sample; A truth value processing module, used to determine the truth values ​​of the first sample and the second sample according to the driving video; A model training module is used to train a trajectory prediction model based on the first sample, the second sample and the true value.

8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.

Citation Information

Cited By

  • Image truth value data generation method and vehicle track prediction model training method

    CN121033155A