Automatic driving track prediction method, device and equipment based on self-distillation learning

The trajectory prediction network is trained through the self-distillation learning method, combined with the target generation module, and the efficiency and accuracy of autonomous driving trajectory prediction in the occlusion scenario are solved, and efficient trajectory prediction and simplified training process are achieved.

CN120470932APending Publication Date: 2025-08-12BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510706978.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing autonomous driving trajectory prediction model has low prediction efficiency and accuracy in occluded traffic scenarios, and the training process is cumbersome and costly.

Method used

The self-distillation learning method is adopted to train the trajectory prediction network through sample complete trajectory data and sample mask trajectory data, and combine the target generation module to achieve end-to-end trajectory prediction, and optimize the network using the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

Benefits of technology

It improves the adaptability and prediction efficiency of trajectory prediction in occluded scenarios, simplifies the training process, and reduces the computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470932A_ABST
    Figure CN120470932A_ABST
Patent Text Reader

Abstract

The invention provides an automatic driving track prediction method, device and equipment based on self-distillation learning, and is applied to the technical field of automatic driving, and the method comprises the steps: firstly obtaining to-be-predicted data, inputting the to-be-predicted data into a track prediction network for processing, and obtaining a track prediction result corresponding to the to-be-predicted data, the trajectory prediction network is obtained by training an initial trajectory prediction network by sample complete trajectory data, sample map data, sample mask trajectory data and a preset loss function; the preset loss function is constructed based on a complete track loss function, a mask track loss function and a maximum mean value difference loss function. According to the technical scheme, the track prediction network obtained by training the complete track data and the mask track data is suitable for track analysis tasks with different complexity degrees, the to-be-predicted data are input into the track prediction network to obtain the track prediction result, automatic processing and prediction of the to-be-predicted track data are achieved, and the track analysis efficiency is improved. And the trajectory prediction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to a method, apparatus, and device for predicting autonomous driving trajectories based on self-distillation learning. Background Art

[0002] In the field of autonomous driving, accurately predicting the movement trajectory of the vehicle itself and surrounding intelligent entities in the next few seconds is an extremely challenging yet crucial task. By accurately grasping the dynamic direction of surrounding vehicles, pedestrians and other intelligent entities, the autonomous driving system can plan the path trajectory in advance and guide the vehicle along a safe and efficient route.

[0003] In the existing technology, a three-stage prediction model of motion prediction, self-supervised learning, and feature distillation is adopted. However, the prediction model cannot adapt to traffic scenes with occlusions when performing trajectory prediction, and the prediction efficiency and accuracy are low.

[0004] Therefore, how to improve the adaptability of autonomous driving trajectory prediction and the efficiency of trajectory prediction has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, and device for autonomous driving trajectory prediction based on self-distillation learning to address the problems of the existing technology in which the prediction model cannot adapt to obstructed traffic scenarios and has low prediction efficiency when performing trajectory prediction.

[0006] In a first aspect, an embodiment of the present application provides an autonomous driving trajectory prediction method based on self-distillation learning, comprising:

[0007] Obtain the data to be predicted;

[0008] The data to be predicted is input into the trajectory prediction network for processing to obtain the trajectory prediction result corresponding to the data to be predicted;

[0009] Among them, the trajectory prediction network is obtained by training the initial trajectory prediction network based on the sample complete trajectory data, sample map data, sample mask trajectory data and the preset loss function; the initial trajectory prediction network includes a target point generation module; the preset loss function is constructed based on the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

[0010] In one possible implementation, the data to be predicted is input into a trajectory prediction network for processing to obtain a trajectory prediction result corresponding to the data to be predicted, including:

[0011] The real-time map data in the data to be predicted is input into the map encoder in the trajectory prediction network for processing to obtain the map encoding features;

[0012] The real-time trajectory data and map coding features in the data to be predicted are input into the agent encoder in the trajectory prediction network for processing to obtain trajectory data features;

[0013] The trajectory data features and map coding features are input into the target generation module in the trajectory prediction network for processing to obtain the trajectory prediction target point;

[0014] The trajectory prediction target point and trajectory data features are input into the trajectory prediction module in the trajectory prediction network for processing to obtain the trajectory prediction result corresponding to the data to be predicted.

[0015] In one possible implementation, the process of training the trajectory prediction network includes:

[0016] Acquire multiple training samples, each training sample including sample complete trajectory data, sample mask trajectory data, and sample map encoding features;

[0017] The sample complete trajectory data, sample mask trajectory data and sample map encoding features are input into the agent encoder in the initial trajectory prediction network for processing to obtain the complete trajectory features and mask trajectory features;

[0018] The complete trajectory features, mask trajectory features, and sample map encoding features are input into the target point generation module in the initial trajectory prediction network for processing to obtain the complete trajectory prediction target points and the mask trajectory prediction target points;

[0019] The complete trajectory features, mask trajectory features, complete trajectory predicted target points, mask trajectory predicted target points and sample map encoding features are input into the trajectory prediction module in the initial trajectory prediction network for processing to obtain the complete trajectory prediction results and mask trajectory prediction results;

[0020] The initial trajectory prediction network is trained according to the complete trajectory prediction results, the mask trajectory prediction results and the preset loss function until the preset number of training times is reached to obtain the trajectory prediction network.

[0021] In one possible implementation, the complete trajectory features, mask trajectory features, complete trajectory predicted target points, mask trajectory predicted target points, and sample map encoding features are input into the trajectory prediction module in the initial trajectory prediction network for processing to obtain complete trajectory prediction results and mask trajectory prediction results, including:

[0022] Establish the correspondence between the preset trajectory query vector and the complete trajectory prediction target point, and the correspondence between the preset trajectory query vector and the mask trajectory prediction target point, and obtain the complete trajectory joint feature and the mask trajectory joint feature;

[0023] The cross-attention module is used to fuse the joint features of the complete trajectory with the sample map encoding features, and the joint features of the masked trajectory with the sample map encoding features to obtain the dynamic features of the target points of the complete trajectory and the dynamic features of the target points of the masked trajectory;

[0024] A multi-head attention mechanism is used to optimize the dynamic features of the complete trajectory target point and the mask trajectory target point respectively, and the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point are obtained;

[0025] The dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the feed-forward neural network FFN module for processing to obtain the complete trajectory prediction results and the mask trajectory prediction results.

[0026] In one possible implementation, the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point are input into the FFN module for processing, including:

[0027] The dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the FFN module for decoding processing. The FFN module generates the complete trajectory prediction results and the mask trajectory prediction results based on the Laplace distribution assumption.

[0028] In a second aspect, an embodiment of the present application provides an autonomous driving trajectory prediction device based on self-distillation learning, comprising:

[0029] An acquisition module is used to obtain data to be predicted;

[0030] The processing module is used to input the data to be predicted into the trajectory prediction network for processing and obtain the trajectory prediction result corresponding to the data to be predicted;

[0031] Among them, the trajectory prediction network is obtained by training the initial trajectory prediction network based on the sample complete trajectory data, sample map data, sample mask trajectory data and the preset loss function; the initial trajectory prediction network includes a target point generation module; the preset loss function is constructed based on the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

[0032] In a possible implementation, the processing module is specifically configured to:

[0033] The real-time map data in the data to be predicted is input into the map encoder in the trajectory prediction network for processing to obtain the map encoding features;

[0034] The real-time trajectory data and map coding features in the data to be predicted are input into the agent encoder in the trajectory prediction network for processing to obtain trajectory data features;

[0035] The trajectory data features and map coding features are input into the target generation module in the trajectory prediction network for processing to obtain the trajectory prediction target point;

[0036] The trajectory prediction target point and trajectory data features are input into the trajectory prediction module in the trajectory prediction network for processing to obtain the trajectory prediction result corresponding to the data to be predicted.

[0037] In one possible implementation, the processing module trains a trajectory prediction network, specifically for:

[0038] Acquire multiple training samples, each training sample including sample complete trajectory data, sample mask trajectory data, and sample map encoding features;

[0039] The sample complete trajectory data, sample mask trajectory data and sample map encoding features are input into the agent encoder in the initial trajectory prediction network for processing to obtain the complete trajectory features and mask trajectory features;

[0040] The complete trajectory features, mask trajectory features, and sample map encoding features are input into the target point generation module in the initial trajectory prediction network for processing to obtain the complete trajectory prediction target points and the mask trajectory prediction target points;

[0041] The complete trajectory features, mask trajectory features, complete trajectory predicted target points, mask trajectory predicted target points and sample map encoding features are input into the trajectory prediction module in the initial trajectory prediction network for processing to obtain the complete trajectory prediction results and mask trajectory prediction results;

[0042] The initial trajectory prediction network is trained according to the complete trajectory prediction results, the mask trajectory prediction results and the preset loss function until the preset number of training times is reached to obtain the trajectory prediction network.

[0043] In a possible implementation, the processing module is specifically configured to:

[0044] Establish the correspondence between the preset trajectory query vector and the complete trajectory prediction target point, and the correspondence between the preset trajectory query vector and the mask trajectory prediction target point, and obtain the complete trajectory joint feature and the mask trajectory joint feature;

[0045] The cross-attention module is used to fuse the joint features of the complete trajectory with the sample map encoding features, and the joint features of the masked trajectory with the sample map encoding features to obtain the dynamic features of the target points of the complete trajectory and the dynamic features of the target points of the masked trajectory;

[0046] A multi-head attention mechanism is used to optimize the dynamic features of the complete trajectory target point and the mask trajectory target point respectively, and the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point are obtained;

[0047] The dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the feed-forward neural network FFN module for processing to obtain the complete trajectory prediction results and the mask trajectory prediction results.

[0048] In a possible implementation, the processing module is specifically configured to:

[0049] The dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the FFN module for decoding processing. The FFN module generates the complete trajectory prediction results and the mask trajectory prediction results based on the Laplace distribution assumption.

[0050] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0051] Memory stores computer-executable instructions;

[0052] The processor executes the computer-executable instructions stored in the memory to implement the method as described in the first aspect or any one of the above-mentioned methods.

[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method of the first aspect or any of the above-mentioned methods.

[0054] In a fifth aspect, an embodiment of the present application provides a computer program, wherein a computer program product includes a computer program, wherein the computer program is stored in a computer-readable storage medium, and at least one processor can read the computer program from the computer-readable storage medium. When at least one processor executes the computer program, the method of the first aspect or any one of the above methods can be implemented.

[0055] The embodiments of the present application provide a method, apparatus, and device for autonomous driving trajectory prediction based on self-distillation learning. The method first obtains the data to be predicted, inputs the data to be predicted into a trajectory prediction network for processing, and obtains a trajectory prediction result corresponding to the data to be predicted. The trajectory prediction network is obtained by training an initial trajectory prediction network using sample complete trajectory data, sample map data, sample masked trajectory data, and a preset loss function; the preset loss function is constructed based on the complete trajectory loss function, the masked trajectory loss function, and the maximum mean difference loss function. This technical solution uses a trajectory prediction network trained with complete trajectory data and masked trajectory data, and is suitable for trajectory analysis tasks of varying complexity. The data to be predicted is input into the trajectory prediction network to obtain a trajectory prediction result, thereby achieving automatic processing and prediction of the trajectory data to be predicted and improving the efficiency of trajectory prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0057] Figure 1 A schematic diagram of the trajectory prediction network architecture provided in an embodiment of the present application;

[0058] Figure 2 A schematic diagram of a flow chart of an autonomous driving trajectory prediction method based on self-distillation learning provided in an embodiment of the present application;

[0059] Figure 3 Schematic diagram of the training process of the trajectory prediction network provided in the embodiment of the present application Figure 1 ;

[0060] Figure 4 Schematic diagram of the training process of the trajectory prediction network provided in the embodiment of the present application Figure 2 ;

[0061] Figure 5 A schematic diagram of the structure of an autonomous driving trajectory prediction device based on self-distillation learning provided in an embodiment of the present application;

[0062] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0063] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0064] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0065] Before introducing the embodiments of the present application, the application background of the embodiments of the present application is first explained:

[0066] In the field of autonomous driving, accurately predicting the movement trajectories of surrounding intelligent agents within the next few seconds is a challenging yet crucial task. By accurately understanding the dynamic movements of surrounding intelligent agents such as vehicles and pedestrians, autonomous driving systems can plan driving paths in advance and guide vehicles along safe and efficient routes. This not only effectively avoids potential collisions but also significantly improves the safety and reliability of autonomous driving. It provides solid support for trajectory planning during the autonomous driving process, driving the development of autonomous driving technology towards smarter and safer directions.

[0067] However, traditional trajectory prediction schemes have many limitations. Most of these schemes are trained on open-source, large-scale motion prediction datasets, which typically assume that the input observations (i.e., coordinate point locations) and output observations of targets (such as vehicles or pedestrians) are complete and accurate. However, in the real world, traffic safety is often seriously threatened by insufficient observations. For example, obstacles, other vehicles, or pedestrians on the road may partially block the target's view, resulting in incomplete or even missing observations. This incomplete observation data can further lead to deviations in trajectory prediction, thereby affecting the decision-making accuracy and safety of the autonomous driving system, and may even cause potential safety hazards.

[0068] Existing techniques for trajectory prediction under partial observation conditions employ multi-stage training. First, a teacher model is constructed in the motion prediction stage. A masking strategy is introduced in the self-supervised learning stage to randomly mask some historical trajectory time steps to simulate scenarios where real observations are missing. A reconstruction branch, similar in structure to the prediction branch, is used to reconstruct the complete historical trajectory based on the partial observations, capturing temporal dependencies. The loss function integrates the motion prediction loss and the reconstruction loss. Finally, in the feature distillation stage, the teacher model parameters are frozen. The student model, whose input is the masked partial observation, is aligned with the encoder and hidden features of the interaction module of the teacher model using mean squared error. This forces the student to mimic the teacher's feature distribution and enhances its predictive power for incomplete data.

[0069] However, the methods adopted in the above prior art have two major disadvantages:

[0070] 1. Performance degradation: During the feature distillation stage, the performance of the distilled student network on complete trajectory data is lower than that of the network trained directly. This is because during the knowledge distillation process, the student model focuses on improving its prediction capabilities for incomplete trajectory data and imitating the features of the teacher model, which to some extent weakens its performance in complete trajectory data scenarios.

[0071] 2. The training process is cumbersome and costly. The entire training process involves multiple stages: first, motion prediction training to obtain a teacher model, then self-supervised learning training, and finally feature distillation training. This multi-stage training process is extremely cumbersome and requires significant investment in time, computing resources, and other costs, significantly increasing the complexity and cost of training.

[0072] Therefore, how to improve the adaptability of autonomous driving trajectory prediction and the efficiency of trajectory prediction has become a technical problem that needs to be solved urgently.

[0073] In response to the technical problems existing in the prior art, the inventors of this application have the following ideas: to address the problem of limited adaptability of autonomous driving trajectory prediction, the initial trajectory prediction network is trained using sample complete trajectory data and sample mask trajectory data, thereby improving the trajectory prediction ability of the trajectory prediction network in obstructed scenarios. To address the problem of low efficiency of autonomous driving trajectory prediction, the data to be predicted is input into the trajectory prediction network for end-to-end processing, and the target generation module is combined to accurately predict the target point of the trajectory, thereby obtaining the trajectory prediction result corresponding to the predicted data, thereby improving the efficiency of trajectory prediction. The trajectory prediction network is obtained by training the initial trajectory prediction network based on sample complete trajectory data, sample map data, sample mask trajectory data and a preset loss function; the initial trajectory prediction network includes a target point generation module; and the preset loss function is constructed based on the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

[0074] Specifically, Figure 1 A schematic diagram of the trajectory prediction network architecture provided in the embodiment of this application is shown as follows: Figure 1 As shown, a brief introduction to the trajectory prediction network involved in the embodiment of the present application is given:

[0075] The trajectory prediction network architecture includes: data input layer, prediction layer, and data output layer.

[0076] The data input layer is used to input complete trajectory data, mask trajectory data, and map data.

[0077] The prediction layer includes an agent encoder, a map encoder, a target point generation module, and a trajectory prediction module. The trajectory prediction module includes a feed-forward neural network (FFN) module, a cross-attention module, etc.

[0078] The data output layer includes outputting the output results of the trajectory prediction module.

[0079] Parts not described in detail are disclosed by the following embodiments.

[0080] The technical solution of the present application is described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0081] It is worth noting that the application fields of the autonomous driving trajectory prediction method, device, electronic device and storage medium based on self-distillation learning adopted in this application are not limited.

[0082] Among them, the execution subject of this application is an electronic device, which can be a server, terminal device, etc.

[0083] Figure 2 A flow chart of the autonomous driving trajectory prediction method based on self-distillation learning provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method may include the following steps:

[0084] Step 21: Obtain the data to be predicted.

[0085] In this step, the autonomous driving system obtains the data to be predicted through various sensors and data interfaces, providing a data basis for the trajectory prediction network to perform trajectory prediction, helping the trajectory prediction network to better understand the current environment and vehicle status.

[0086] The data to be predicted includes the autonomous vehicle's real-time position, speed, acceleration, steering angle, and data from environmental perception modules (such as cameras, radar, and lidar). Furthermore, the autonomous driving system can obtain real-time traffic conditions, road construction information, and other environmental data that affects vehicle operation through a traffic information interface.

[0087] Step 22: Input the data to be predicted into the trajectory prediction network for processing to obtain the trajectory prediction result corresponding to the data to be predicted.

[0088] Among them, the trajectory prediction network is obtained by training the initial trajectory prediction network based on the sample complete trajectory data, sample map data, sample mask trajectory data and the preset loss function; the initial trajectory prediction network includes a target point generation module; the preset loss function is constructed based on the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

[0089] In this step, the data input layer transmits the data to be predicted to the trained trajectory prediction network. The trajectory prediction network predicts the future trajectory by analyzing the trajectory data (i.e., vehicle dynamic information) and map data (such as information about the vehicle's surrounding environment) in the data to be predicted. The obtained future trajectory prediction result provides a reliable basis for the vehicle's subsequent path planning, decision-making, and control.

[0090] Among them, the trajectory prediction results include the expected position, speed, acceleration and possible turning trajectory of the autonomous driving vehicle in the future.

[0091] Optionally, step 22 may be implemented as follows:

[0092] Step 1: Input the real-time map data in the data to be predicted into the map encoder in the trajectory prediction network for processing to obtain map encoding features.

[0093] In this implementation, the real-time map data in the data to be predicted is input into the map encoder in the trajectory prediction network for processing. By analyzing the real-time map data, the geographical features and road information of the current environment, such as road type, intersection, road condition, and traffic signal, are extracted to obtain the map coding features, which provide spatial constraints and guidance for subsequent trajectory prediction.

[0094] Through the above implementation, the trajectory prediction network can make corresponding trajectory adjustments according to different map environments, ensuring that the prediction results are consistent with actual road conditions and traffic regulations.

[0095] Step 2: Input the real-time trajectory data and map coding features in the data to be predicted into the agent encoder in the trajectory prediction network for processing to obtain trajectory data features.

[0096] In this implementation, the real-time trajectory data and map coding features in the data to be predicted are input into the intelligent agent encoder in the trajectory prediction network for processing. The intelligent agent encoder extracts dynamic features such as speed, acceleration, and steering angle from the vehicle's real-time trajectory data, and then jointly embeds the obtained dynamic features with the map coding features for representation. This enables the intelligent agent encoder to more accurately understand the vehicle's driving behavior in the current map environment and obtain trajectory data features.

[0097] Step 3: Input the trajectory data features and map coding features into the target generation module in the trajectory prediction network for processing to obtain the trajectory prediction target point.

[0098] In this implementation, trajectory data features and map coding features are input into the target generation module in the trajectory prediction network for processing. Based on the input trajectory data and map coding features, the possible target positions of the vehicle at key time nodes within the next few time steps are predicted, and the trajectory prediction target points are obtained. This can provide a clear navigation target for the trajectory prediction module, making the predicted trajectory more in line with actual driving needs.

[0099] Among them, the trajectory prediction target point obtained by the target generation module is generated based on the current driving trajectory and map environment information, which can effectively reflect the possible driving path of the autonomous driving vehicle.

[0100] Step 4: Input the trajectory prediction target point and trajectory data features into the trajectory prediction module in the trajectory prediction network for processing to obtain the trajectory prediction result corresponding to the data to be predicted.

[0101] In this implementation, the trajectory prediction target points and trajectory data features are input into the trajectory prediction module in the trajectory prediction network for processing. Based on the trajectory prediction target points provided by the target generation module and the previously extracted trajectory data features, the current vehicle status and historical behavior are analyzed. Combined with the map features, the vehicle's trajectory in the future is predicted, which can accurately reflect the vehicle's movement trend and help the subsequent decision-making module of the autonomous driving system make reasonable driving decisions, ensuring that the autonomous driving system can operate safely and effectively in complex traffic environments.

[0102] An embodiment of the present application provides an autonomous driving trajectory prediction method based on self-distillation learning. The method first obtains the data to be predicted, inputs the data to be predicted into a trajectory prediction network for processing, and obtains a trajectory prediction result corresponding to the data to be predicted. The trajectory prediction network is obtained by training an initial trajectory prediction network using sample complete trajectory data, sample map data, sample masked trajectory data, and a preset loss function; the preset loss function is constructed based on the complete trajectory loss function, the masked trajectory loss function, and the maximum mean difference loss function. This technical solution uses a trajectory prediction network trained with complete trajectory data and masked trajectory data, and is suitable for trajectory analysis tasks of varying complexity. The data to be predicted is input into the trajectory prediction network to obtain a trajectory prediction result, achieving automatic processing and prediction of the trajectory data to be predicted and improving the efficiency of trajectory prediction.

[0103] Based on the above embodiments, Figure 3 Schematic diagram of the training process of the trajectory prediction network provided in the embodiment of the present application Figure 1 ,like Figure 3As shown, the training process may include the following steps:

[0104] Step 31: Obtain multiple training samples.

[0105] Each training sample includes sample complete trajectory data, sample mask trajectory data, and sample map encoding features.

[0106] In this step, multiple training samples are obtained from historical data. Each sample contains complete trajectory data, masked trajectory data, and corresponding sample map encoding features, which are used to train the initial trajectory prediction network to help the initial trajectory prediction network learn diverse environment and trajectory features.

[0107] Among them, multiple training samples cover the driving conditions of vehicles in different road environments, as well as corresponding environmental features such as road types, traffic signals, etc.

[0108] Complete trajectory data refers to the actual driving trajectory of a vehicle over a period of time, while masked trajectory data refers to the complete trajectory data after masking. Using both complete trajectory data and masked trajectory data as training samples helps improve the robustness of the trajectory prediction network for driving in traffic scenarios with occlusions.

[0109] In addition, the sample map coding feature is a feature vector obtained after encoding the real-time map data, which reflects the detailed geographic information of the area where the autonomous driving vehicle is located.

[0110] Optionally, step 31 may be implemented as follows:

[0111] Step 1: Obtain the vehicle's complete sample trajectory data and corresponding sample map data.

[0112] Among them, the sample complete trajectory data is the historical trajectory data of the vehicle.

[0113] In this implementation, the complete trajectory data of the autonomous driving vehicle is obtained from the historical data, and the map data corresponding to the complete trajectory data is also obtained.

[0114] Among them, the complete trajectory data of the autonomous driving vehicle represents the vehicle's driving path over the past period of time, including the complete trajectory data and the corresponding status information of surrounding intelligent entities (that is, surrounding vehicles other than the autonomous driving vehicle itself). The map data corresponding to the complete trajectory data represents the geographic information of the area where the vehicle is traveling, such as road type, traffic signals and other map-related information.

[0115] Step 2: Perform random masking on the complete trajectory data of the sample to obtain the masked trajectory data of the sample.

[0116] In this implementation, the mask module in the data input layer (such as Figure 1 The complete trajectory data obtained is randomly masked (as shown in the figure). The random masking technology is used to simulate the loss or noise of the complete trajectory data, which helps the initial trajectory prediction network to process incomplete or noisy data during training and improves the prediction robustness of the initial trajectory prediction network for trajectory data of different qualities.

[0117] Among them, mask processing will delete certain trajectory points in the complete trajectory data and generate incomplete trajectory data as mask trajectory data. The mask trajectory data includes trajectory data and corresponding surrounding intelligent agents (that is, surrounding vehicles other than the autonomous driving vehicle itself) status information.

[0118] Step 3: Input the sample map data into the map encoder in the initial trajectory prediction network for processing to obtain the sample map encoding features.

[0119] In this implementation, the sample map data is input into the map encoder in the initial trajectory prediction network for processing, and the sample map data is converted into sample map encoding features.

[0120] Among them, the sample map data includes detailed geographic information such as lane line information and drivable areas in the current map area, which can help the trajectory prediction network better understand the environment in which the autonomous driving vehicle is located and its impact on the driving path.

[0121] Step 32: Input the sample complete trajectory data, the sample mask trajectory data, and the sample map encoding features into the agent encoder in the initial trajectory prediction network for processing to obtain the complete trajectory features and the mask trajectory features.

[0122] In this step, the complete trajectory data, masked trajectory data, and sample map encoding features are input into the intelligent encoder in the initial trajectory prediction network for processing. The intelligent encoder extracts dynamic features such as motion state, speed, acceleration, steering angle, etc. from the complete trajectory data of the autonomous driving vehicle and combines them with the map encoding features for joint embedding representation. The dynamic features in the masked trajectory data are jointly embedded with the map encoding features to obtain the complete trajectory features and masked trajectory features, respectively.

[0123] Among them, the complete trajectory features and mask trajectory features obtained by the intelligent encoder are processed by the maximum mean difference loss function, ensuring that the distribution of the mask trajectory data prediction is consistent with the result distribution of the complete trajectory prediction, thereby improving the generalization ability of the initial trajectory prediction network in scenes with occlusion and improving the environmental applicability of the trajectory prediction model.

[0124] In particular, the masked trajectory data provides the trajectory prediction network with training data for incomplete or partially missing data, helping the trajectory prediction network improve its ability to process unknown data in complex environments.

[0125] In one possible implementation, the complete trajectory feature is the current position coordinates of the autonomous driving vehicle, the lane line in which the vehicle is located, and the associated feature information of the adjacent lane lines.

[0126] In addition, the complete trajectory feature F can be expressed by the mathematical formula as follows:

[0127] F={F A ,F X ,F X→A}

[0128] Where, F A is the feature of the surrounding intelligent agent (i.e., the target around the autonomous vehicle) in the complete trajectory data, F X is the agent feature in the complete trajectory data (i.e., the autonomous driving vehicle’s own goal), F X→A It is the interaction feature (such as geometric position relationship) between the agent and surrounding agents in the complete trajectory data.

[0129] Mask trajectory feature F P It can be expressed mathematically as follows:

[0130]

[0131] Where, is the surrounding agent features in the mask trajectory data, is the agent feature in the mask trajectory data, is the interaction feature between the agent and surrounding agents in the mask trajectory data.

[0132] Step 33: Input the complete trajectory features, mask trajectory features, and sample map encoding features into the target point generation module in the initial trajectory prediction network for processing to obtain the complete trajectory predicted target points and the mask trajectory predicted target points.

[0133] In this step, the complete trajectory features, mask trajectory features, and sample map encoding features are input into the target point generation module in the initial trajectory prediction network for processing. The target point generation module predicts the possible target position of the autonomous driving vehicle at key time points in the future under the complete trajectory by combining the complete trajectory features and the map encoding features, and obtains the predicted target point of the complete trajectory. At the same time, the module predicts the possible target position of the autonomous driving vehicle at key time points in the future under the mask trajectory by combining the mask trajectory features and the map encoding features, and obtains the predicted target point of the mask trajectory.

[0134] Among them, the complete trajectory prediction target point σ t , and the target point feature τ corresponding to the complete trajectory prediction target point can be expressed by the mathematical formula as follows:

[0135] τ,σ t =f target (F A ,F X ,F X→A ,F M )

[0136] Where, F M is the map encoding feature, f target () is the target point generating function.

[0137] Mask trajectory prediction target point And the target point feature τ corresponding to the mask trajectory prediction target point p It can be expressed mathematically as follows:

[0138]

[0139] Where, F M Encode features for the map.

[0140] Step 34: Input the complete trajectory features, mask trajectory features, complete trajectory predicted target points, mask trajectory predicted target points, and sample map encoding features into the trajectory prediction module in the initial trajectory prediction network for processing to obtain complete trajectory prediction results and mask trajectory prediction results.

[0141] In this step, the complete trajectory features, mask trajectory features, complete trajectory predicted target points, mask trajectory predicted target points, and sample map encoding features are input into the trajectory prediction module in the initial trajectory prediction network to analyze the current driving state, historical trajectory, and driving behavior of the autonomous driving vehicle. By combining the complete trajectory predicted target points with the complete trajectory features, the vehicle's driving path in the future several moments under the complete trajectory is predicted to obtain a complete trajectory prediction result. At the same time, by combining the mask trajectory predicted target points with the mask trajectory features, the vehicle's driving path in the future several moments under the mask trajectory is predicted to obtain a mask trajectory prediction result. The obtained prediction result will reflect the movement trend of the autonomous driving vehicle and help the autonomous driving system make further decisions and path planning.

[0142] Step 35: Train the initial trajectory prediction network according to the complete trajectory prediction result, the mask trajectory prediction result, and the preset loss function until the preset number of training times is reached to obtain the trajectory prediction network.

[0143] In this step, during the model training process, the error between the trajectory prediction result and the actual trajectory is compared, and the error is back-propagated to the initial trajectory prediction network for iterative training. During the iterative training process, the parameters of the initial trajectory prediction network are continuously optimized based on the preset loss function until the preset number of training times is reached to obtain the trajectory prediction network.

[0144] This process not only improves the prediction accuracy of the trajectory prediction network, but also enhances the network's adaptability to complex traffic scenarios and diverse data. The trained trajectory prediction network has strong generalization capabilities and can handle trajectory prediction tasks in different driving environments and complex traffic conditions, providing more reliable path planning support for autonomous driving systems.

[0145] Among them, the preset loss function It is constructed based on the complete trajectory loss function, the mask trajectory loss function, and the maximum mean difference loss function, which can be expressed mathematically as follows:

[0146]

[0147] Where, is the loss function for the complete trajectory prediction branch, is the loss function of the mask trajectory prediction branch, is the maximum mean difference loss.

[0148] By minimizing the feature distribution distance in the feature space, the latent space alignment of the complete trajectory feature and the mask trajectory feature is achieved, which can be expressed by the following mathematical formula:

[0149]

[0150] In the formula, n represents the number of samples, F Xi represents the agent features in the complete trajectory data of the i-th sample, F Ai represents the surrounding agent features in the mask trajectory data of the i-th sample, is the interaction feature between the agent and surrounding agents in the complete trajectory data of the i-th sample, represents the agent features in the mask trajectory data of the i-th sample, represents the surrounding agent features in the mask trajectory data of the i-th sample, is the interaction feature between the agent and surrounding agents in the mask trajectory data of the i-th sample.

[0151] The loss function of the complete trajectory prediction branch above is It consists of three parts and can be expressed mathematically as follows:

[0152]

[0153] Where, represents the complete trajectory prediction target point prediction loss, is the regression loss of the initial trajectory prediction network, is the classification loss of the initial trajectory prediction network. Similarly, the loss function of the mask trajectory prediction branch is It can be expressed as the following mathematical formula:

[0154]

[0155] Where, Denotes the target point prediction loss of mask trajectory prediction.

[0156] The embodiment of the present application provides an autonomous driving trajectory prediction method based on self-distillation learning. The method first obtains multiple training samples, then inputs the sample complete trajectory data, the sample mask trajectory data and the sample map encoding features into the intelligent agent encoder in the initial trajectory prediction network for processing, and obtains the complete trajectory features and the mask trajectory features, and inputs the complete trajectory features, the mask trajectory features and the sample map encoding features into the target point generation module in the initial trajectory prediction network for processing, and obtains the complete trajectory predicted target point and the mask trajectory predicted target point, and inputs the complete trajectory features, the mask trajectory features, the complete trajectory predicted target point, the mask trajectory predicted target point and the sample map encoding features into the trajectory prediction module in the initial trajectory prediction network for processing, and obtains the complete trajectory. The trajectory prediction results and mask trajectory prediction results are obtained. Finally, the initial trajectory prediction network is trained according to the complete trajectory prediction results, the mask trajectory prediction results and the preset loss function until the preset training times are reached to obtain the trajectory prediction network. This technical solution first obtains multiple training samples and inputs the sample data into different modules of the initial trajectory prediction network for processing, and finally obtains the complete trajectory prediction results and the mask trajectory prediction results. Finally, the network is iteratively trained in combination with the preset loss function until the preset training times are reached, thereby obtaining the optimized trajectory prediction network. Through a multi-stage processing flow, the features of the complete trajectory and the mask trajectory are gradually extracted and optimized to generate high-quality trajectory prediction results, which improves the prediction efficiency while also enhancing the accuracy and robustness of the prediction.

[0157] Based on the above embodiments, Figure 4 Schematic diagram of the training process of the trajectory prediction network provided in the embodiment of the present application Figure 2 ,like Figure 4 As shown, step 34 may include the following implementation steps:

[0158] Step 41: Establish a correspondence between a preset trajectory query vector and a target point predicted by the complete trajectory, and a correspondence between a preset trajectory query vector and a target point predicted by the mask trajectory, and obtain a complete trajectory joint feature and a mask trajectory joint feature.

[0159] In this step, a joint feature vector is established for the trajectory points in the trajectory query vector and the complete trajectory prediction target points and the mask trajectory prediction target points, respectively, to obtain the complete trajectory joint feature and the mask trajectory joint feature.

[0160] The preset trajectory query vector may be adjusted according to actual conditions and is not limited here.

[0161] In a possible implementation, there are 5 preset trajectory query vectors.

[0162] Step 42: Use the cross-attention module to fuse the complete trajectory joint features with the sample map encoding features, and the mask trajectory joint features with the sample map encoding features to obtain the dynamic features of the complete trajectory target point and the dynamic features of the mask trajectory target point.

[0163] In this step, the complete trajectory joint feature uses the attention mechanism to filter relevant elements in the map encoding features (such as adjacent lane lines, traffic signs, etc.) to generate the dynamic features of the target point of the complete trajectory. The masked trajectory joint feature first compensates for the missing information based on the map encoding features (for example, inferring the target point in the occluded area through the lane line curvature) to obtain the dynamic features of the target point of the masked trajectory.

[0164] In one possible implementation, the dynamic features of the complete trajectory target point can be the geometric constraint information features and traffic light constraint information features of the autonomous driving vehicle in the left turn lane, and the dynamic features of the masked trajectory target point can be the target point information features that are obscured when the lane line continuity is blocked by buildings.

[0165] Step 43: Use a multi-head attention mechanism to optimize the dynamic features of the complete trajectory target point and the dynamic features of the mask trajectory target point respectively, to obtain the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point.

[0166] In this step, the dynamic features of the complete trajectory target points and the masked trajectory target points are split into multiple feature subspaces respectively. The multi-head attention mechanism is used to focus on multiple different feature subspaces in parallel, so that the trajectory prediction network can perform specific analysis on the dynamic features of the complete trajectory target points and the masked trajectory target points from multiple angles, strengthen the trajectory prediction network's understanding of the dynamic features of the complete trajectory target points and the masked trajectory target points, and improve the trajectory prediction network's adaptability to different traffic environments.

[0167] Among them, the multi-head attention mechanism can be expressed by the mathematical formula as follows:

[0168]

[0169] Where σ t is the embedded feature of the target sequence, which represents the main information to be predicted (i.e., the dynamic features of the target points of the complete trajectory and the dynamic features of the target points of the mask trajectory). q W is the embedded feature of the query sequence, which represents the contextual state information features at the current moment, such as the fusion representation of the vehicle's real-time motion features (such as speed and heading angle) and the environmental interaction features. Q represents a set of query vectors obtained by linear transformation of the input features, K represents the key vector obtained by linear transformation of the input features, and V represents the value vector obtained by linear transformation of the input features. When the similarity between Q and K is high, the corresponding V will be given a higher weight. q , W k , W v Represent three independent weight matrices respectively.

[0170] Among them, in the attention mechanism, σ t and σ q Can be achieved through W q and W k These two trainable weight matrices are jointly transformed to transform σ t +σ q Mapped to multidimensional query space and key space respectively, W q Generate query vector Q, W k Generate the key vector K. This design makes the information of the target sequence and the query sequence form a dynamic coupling in the feature space. For example, when a vehicle approaches an intersection, σ t (e.g., preset steering path) and σ q The superposition of real-time steering wheel angle can strengthen the attention weight of steering intention.

[0171] In addition, W v Acting alone on σ q Generating a value vector V allows the attention mechanism to focus more on the real-time state of the query sequence. For example, in trajectory prediction, V primarily carries the dynamic obstacle features perceived by the sensor at the current moment, while K retains static environmental constraints such as lane lines.

[0172] Step 44: Input the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point into the FFN module for processing to obtain the complete trajectory prediction result and the mask trajectory prediction result.

[0173] In this step, the dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the FFN module for processing. Through the hierarchical structure of the neural network, the input dynamic features are further analyzed and transformed to generate the final complete trajectory prediction results and mask trajectory prediction results.

[0174] Optionally, step 44 may be implemented as follows:

[0175] The dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the FFN module for decoding processing. The FFN module generates the complete trajectory prediction results and the mask trajectory prediction results based on the Laplace distribution assumption.

[0176] In this implementation, the FFN module assumes that the dynamic features of the input optimized complete trajectory target points and the dynamic features of the optimized mask trajectory target points are Laplace distributed, and then effectively decodes the input dynamic features and outputs accurate and stable trajectory prediction results.

[0177] Among them, the Laplace distribution can be expressed by the following mathematical formula:

[0178] P(X pred |F,F M ,σ t )

[0179] Where, X pred is the set of spatial coordinates of the predicted trajectory.

[0180] The complete trajectory prediction result can be expressed by the following mathematical formula:

[0181] X pred =f pred (F,F M ,σ t )

[0182] Where, f pred () is the trajectory prediction function.

[0183] The mask trajectory prediction result can be expressed by the following mathematical formula:

[0184] X pred =f pred (F p ,F M ,σ t )

[0185] The embodiment of the present application provides an autonomous driving trajectory prediction method based on self-distillation learning. The method first establishes a correspondence between a preset trajectory query vector and a complete trajectory prediction target point, and a correspondence between a preset trajectory query vector and a mask trajectory prediction target point, to obtain a complete trajectory joint feature and a mask trajectory joint feature. Then, a cross-attention module is used to fuse the complete trajectory joint feature with the sample map encoding feature, and the mask trajectory joint feature with the sample map encoding feature, to obtain the dynamic features of the complete trajectory target point and the dynamic features of the mask trajectory target point. A multi-head attention mechanism is used to optimize the dynamic features of the complete trajectory target point and the dynamic features of the mask trajectory target point respectively, to obtain the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point. Finally, the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point are input into the FFN module for processing to obtain the complete trajectory prediction result and the mask trajectory prediction result. This technical solution optimizes the dynamic features of the complete trajectory target point and the mask trajectory target point through the cross-attention module and the multi-head attention mechanism. The FFN module enhances the feature processing capability and improves the accuracy and robustness of trajectory prediction.

[0186] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0187] Figure 5 This is a schematic diagram of the structure of the autonomous driving trajectory prediction device based on self-distillation learning provided in the embodiment of the present application. Figure 5 As shown, the device includes:

[0188] An acquisition module 51 is used to acquire data to be predicted;

[0189] The processing module 52 is used to input the data to be predicted into the trajectory prediction network for processing to obtain the trajectory prediction result corresponding to the data to be predicted;

[0190] Among them, the trajectory prediction network is obtained by training the initial trajectory prediction network based on the sample complete trajectory data, sample map data, sample mask trajectory data and the preset loss function; the initial trajectory prediction network includes a target point generation module; the preset loss function is constructed based on the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

[0191] In a possible implementation, the processing module 52 is specifically configured to:

[0192] The real-time map data in the data to be predicted is input into the map encoder in the trajectory prediction network for processing to obtain the map encoding features;

[0193] The real-time trajectory data and map coding features in the data to be predicted are input into the agent encoder in the trajectory prediction network for processing to obtain trajectory data features;

[0194] The trajectory data features and map coding features are input into the target generation module in the trajectory prediction network for processing to obtain the trajectory prediction target point;

[0195] The trajectory prediction target point and trajectory data features are input into the trajectory prediction module in the trajectory prediction network for processing to obtain the trajectory prediction result corresponding to the data to be predicted.

[0196] In one possible implementation, the processing module 52 trains a trajectory prediction network, specifically for:

[0197] Acquire multiple training samples, each training sample including sample complete trajectory data, sample mask trajectory data, and sample map encoding features;

[0198] The sample complete trajectory data, sample mask trajectory data and sample map encoding features are input into the agent encoder in the initial trajectory prediction network for processing to obtain the complete trajectory features and mask trajectory features;

[0199] The complete trajectory features, mask trajectory features, and sample map encoding features are input into the target point generation module in the initial trajectory prediction network for processing to obtain the complete trajectory prediction target points and the mask trajectory prediction target points;

[0200] The complete trajectory features, mask trajectory features, complete trajectory predicted target points, mask trajectory predicted target points and sample map encoding features are input into the trajectory prediction module in the initial trajectory prediction network for processing to obtain the complete trajectory prediction results and mask trajectory prediction results;

[0201] The initial trajectory prediction network is trained according to the complete trajectory prediction results, the mask trajectory prediction results and the preset loss function until the preset number of training times is reached to obtain the trajectory prediction network.

[0202] In a possible implementation, the processing module 52 is specifically configured to:

[0203] Establish the correspondence between the preset trajectory query vector and the complete trajectory prediction target point, and the correspondence between the preset trajectory query vector and the mask trajectory prediction target point, and obtain the complete trajectory joint feature and the mask trajectory joint feature;

[0204] The cross-attention module is used to fuse the joint features of the complete trajectory with the sample map encoding features, and the joint features of the masked trajectory with the sample map encoding features to obtain the dynamic features of the target points of the complete trajectory and the dynamic features of the target points of the masked trajectory;

[0205] A multi-head attention mechanism is used to optimize the dynamic features of the complete trajectory target point and the mask trajectory target point respectively, and the optimized dynamic features of the complete trajectory target point and the optimized dynamic features of the mask trajectory target point are obtained;

[0206] The dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the feed-forward neural network FFN module for processing to obtain the complete trajectory prediction results and the mask trajectory prediction results.

[0207] In a possible implementation, the processing module 52 is specifically configured to:

[0208] The dynamic features of the optimized complete trajectory target points and the optimized mask trajectory target points are input into the FFN module for decoding processing. The FFN module generates the complete trajectory prediction results and the mask trajectory prediction results based on the Laplace distribution assumption.

[0209] It should be noted that it should be understood that the division of the various modules of the above device is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. Moreover, these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. In addition, these modules can be fully or partially integrated together or implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the hardware integrated logic circuit in the processor element or by instructions in the form of software.

[0210] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the electronic device may include: a processor 61, a memory 62, and computer program instructions stored in the memory 62 and executable on the processor 61. When the processor 61 executes the computer program instructions, the method provided in any of the aforementioned embodiments is implemented.

[0211] Optionally, the above-mentioned components of the electronic device may be connected via a system bus.

[0212] The memory 62 may be a separate storage unit or a storage unit integrated in the processor 61. The number of the processor 61 may be one or more.

[0213] It should be understood that the processor 61 can be a central processing unit (CPU), or other general-purpose processors 61, digital signal processors 61 (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor 61 can be a microprocessor 61 or any conventional processor 61. The steps of the method disclosed in this application can be directly implemented by the hardware processor 61 or performed by a combination of hardware and software modules in the processor 61.

[0214] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. A system bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, the figure uses only one thick line, but this does not imply that there is only one bus or only one type of bus. Memory 62 may include random access memory 62 (RAM) and may also include non-volatile memory 62 (NVM), such as at least one disk storage 62.

[0215] All or part of the steps of the above-mentioned method embodiments can be completed by hardware associated with program instructions. The aforementioned program can be stored in a readable memory 62. When executed, the program performs the steps of the above-mentioned method embodiments; and the aforementioned memory 62 (storage medium) includes: read-only memory 62 (ROM), RAM, flash memory 62, hard disk, solid-state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.

[0216] The electronic device provided in the embodiments of the present application can be used to execute the method provided in any of the above method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.

[0217] An embodiment of the present application provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed on a computer, the computer executes the above method.

[0218] The computer-readable storage medium mentioned above may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0219] Optionally, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0220] An embodiment of the present application also provides a computer program product, which includes a computer program. The computer program is stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, the above method can be implemented.

[0221] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for autonomous driving trajectory prediction based on self-distillation learning, characterized in that: include: Obtain the data to be predicted; Inputting the data to be predicted into a trajectory prediction network for processing to obtain a trajectory prediction result corresponding to the data to be predicted; Among them, the trajectory prediction network is obtained by training the initial trajectory prediction network based on sample complete trajectory data, sample map data, sample mask trajectory data and a preset loss function; the initial trajectory prediction network includes a target point generation module; the preset loss function is constructed based on the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

2. The method according to claim 1, characterized in that The step of inputting the data to be predicted into a trajectory prediction network for processing to obtain a trajectory prediction result corresponding to the data to be predicted includes: Inputting the real-time map data in the data to be predicted into a map encoder in the trajectory prediction network for processing to obtain map encoding features; Inputting the real-time trajectory data and the map coding features in the data to be predicted into the agent encoder in the trajectory prediction network for processing to obtain trajectory data features; Inputting the trajectory data features and the map coding features into the target generation module in the trajectory prediction network for processing to obtain a trajectory prediction target point; The trajectory prediction target point and the trajectory data features are input into the trajectory prediction module in the trajectory prediction network for processing to obtain a trajectory prediction result corresponding to the data to be predicted.

3. The method according to claim 1, characterized in that The process of training the trajectory prediction network includes: Acquire a plurality of training samples, each of the training samples including the sample complete trajectory data, the sample mask trajectory data, and a sample map encoding feature; Inputting the sample complete trajectory data, the sample mask trajectory data and the sample map encoding features into the agent encoder in the initial trajectory prediction network for processing to obtain complete trajectory features and mask trajectory features; Inputting the complete trajectory features, the mask trajectory features, and the sample map encoding features into the target point generation module in the initial trajectory prediction network for processing to obtain the complete trajectory prediction target point and the mask trajectory prediction target point; Inputting the complete trajectory features, the mask trajectory features, the complete trajectory predicted target points, the mask trajectory predicted target points, and the sample map encoding features into the trajectory prediction module in the initial trajectory prediction network for processing to obtain a complete trajectory prediction result and a mask trajectory prediction result; The initial trajectory prediction network is trained according to the complete trajectory prediction result, the mask trajectory prediction result, and the preset loss function until a preset number of training times is reached to obtain the trajectory prediction network.

4. The method according to claim 3, characterized in that The obtaining of multiple training samples includes: Obtaining sample complete trajectory data of the vehicle and corresponding sample map data, wherein the sample complete trajectory data is historical trajectory data of the vehicle; Performing random masking on the sample complete trajectory data to obtain the sample masked trajectory data; The sample map data is input into the map encoder in the initial trajectory prediction network for processing to obtain the sample map encoding features.

5. The method according to claim 3, characterized in that The complete trajectory feature, the mask trajectory feature, the complete trajectory predicted target point, the mask trajectory predicted target point, and the sample map encoding feature are input into the trajectory prediction module in the initial trajectory prediction network for processing to obtain a complete trajectory prediction result and a mask trajectory prediction result, including: Establishing a correspondence between a preset trajectory query vector and the complete trajectory prediction target point, and a correspondence between a preset trajectory query vector and the mask trajectory prediction target point, to obtain a complete trajectory joint feature and a mask trajectory joint feature; A cross-attention module is used to fuse the complete trajectory joint feature with the sample map encoding feature, and the mask trajectory joint feature with the sample map encoding feature to obtain the dynamic features of the complete trajectory target point and the dynamic features of the mask trajectory target point; A multi-head attention mechanism is used to respectively optimize the dynamic features of the complete trajectory target point and the dynamic features of the mask trajectory target point, thereby obtaining optimized dynamic features of the complete trajectory target point and optimized dynamic features of the mask trajectory target point; The dynamic features of the optimized complete trajectory target point and the dynamic features of the optimized mask trajectory target point are input into the feed-forward neural network FFN module for processing to obtain the complete trajectory prediction result and the mask trajectory prediction result.

6. The method according to claim 5, characterized in that The step of inputting the dynamic features of the optimized complete trajectory target point and the dynamic features of the optimized mask trajectory target point into the FFN module for processing includes: The dynamic features of the optimized complete trajectory target point and the dynamic features of the optimized mask trajectory target point are input into the FFN module for decoding processing. The FFN module generates the complete trajectory prediction result and the mask trajectory prediction result based on the Laplace distribution assumption.

7. An autonomous driving trajectory prediction device based on self-distillation learning, characterized in that: include: An acquisition module is used to obtain data to be predicted; A processing module, configured to input the data to be predicted into a trajectory prediction network for processing, and obtain a trajectory prediction result corresponding to the data to be predicted; Among them, the trajectory prediction network is obtained by training the initial trajectory prediction network based on sample complete trajectory data, sample map data, sample mask trajectory data and a preset loss function; the preset loss function is constructed based on the complete trajectory loss function, the mask trajectory loss function and the maximum mean difference loss function.

8. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program, characterized in that The computer program includes a computer program, which is stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when the at least one processor executes the computer program, the method described in any one of claims 1 to 6 can be implemented.