Trajectory prediction method and device and vehicle

By acquiring and fusing the position sequence characteristics and lane segment sequence characteristics in vehicle trajectory prediction, and using attention mechanism, the problem of inaccurate trajectory prediction in the prior art is solved, achieving higher prediction accuracy and robustness.

CN120003530APending Publication Date: 2025-05-16CHONGQING CHANGAN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510095366.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has the problem of inaccurate prediction of automobile trajectory, especially because the sparseness of traffic scene information cannot be effectively utilized.

Method used

By obtaining the position sequence characteristics of the vehicle to be predicted and the lane segment sequence characteristics, and using attention mechanism calculations, multi-dimensional information is fused to improve prediction accuracy.

Benefits of technology

It improves the accuracy and robustness of trajectory prediction, can more accurately understand the vehicle behavior patterns in complex traffic scenarios, and generate more reliable predicted trajectories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120003530A_ABST
    Figure CN120003530A_ABST
Patent Text Reader

Abstract

The invention relates to a trajectory prediction method and device and a vehicle, and relates to the technical field of automobiles, in particular to the technical field of automobile trajectory prediction. The method comprises the following steps: acquiring a first position sequence feature and a lane segment sequence feature of a to-be-predicted vehicle; the first position sequence feature is used for representing position features of the to-be-predicted vehicle at multiple historical moments; the lane segment sequence features are used for representing the position relationship between the to-be-predicted vehicle and surrounding lane segments; performing attention mechanism operation based on the first position sequence feature and the lane segment sequence feature to obtain a first lane segment fusion feature; and obtaining a prediction track of the to-be-predicted vehicle based on the first lane segment fusion feature. Therefore, the accuracy of track prediction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automobile technology, in particular to the field of automobile trajectory prediction technology, and specifically to a trajectory prediction method, device and vehicle. Background Art

[0002] As the leader in the transformation and upgrading of the vehicle industry, autonomous vehicles are highly valued and recognized worldwide. The core of this technology is to enable vehicles to make accurate and timely decisions in various complex traffic environments to ensure driving safety and efficiency. In order to achieve this goal, it is particularly important to predict the intentions or movement trajectories of surrounding vehicles (such as other vehicles, pedestrians, etc.) in advance.

[0003] In the related technology, the proposed Multi-PPTP uses a Mobile Net lightweight network as a feature encoder for the raster image, but it cannot make good use of the sparsity of traffic scene information, resulting in inaccurate prediction of trajectories.

[0004] In the related art, a vehicle trajectory feature extraction method is proposed to provide standardized feature data for vehicle trajectory classification or prediction. However, there is a problem that the introduction of sampling increases the uncertainty of the encoding process, resulting in inaccurate prediction of the trajectory. Summary of the invention

[0005] The present application provides a trajectory prediction method, device and vehicle to at least solve the technical problem of low feasibility of trajectory prediction in the related art. The technical solution of the present application is as follows:

[0006] According to a first aspect provided by the present application, a trajectory prediction method is provided, the method comprising: obtaining first position sequence features and lane segment sequence features of a vehicle to be predicted; the first position sequence features are used to characterize the position features of the vehicle to be predicted at multiple historical moments; the lane segment sequence features are used to characterize the positional relationship between the vehicle to be predicted and the surrounding lane segments; performing an attention mechanism operation based on the first position sequence features and the lane segment sequence features to obtain a first lane segment fusion feature; and obtaining a predicted trajectory of the vehicle to be predicted based on the first lane segment fusion feature.

[0007] Based on the above technical solution, the embodiment of the present application integrates multi-dimensional information (first position sequence features and lane segment sequence features) and, through the attention mechanism operation, enables the model to automatically learn and emphasize the position sequence features and lane segment sequence features that have an important impact on the prediction results, thereby improving the accuracy and robustness of the predicted trajectory.

[0008] In some embodiments, based on the first lane segment fusion feature, a predicted trajectory of the vehicle to be predicted is obtained, including: obtaining vehicle interaction features and second lane segment fusion features corresponding to the surrounding vehicles of the vehicle to be predicted; the vehicle interaction features are used to characterize the positional relationship between the vehicle to be predicted and the surrounding vehicles; based on the first lane segment fusion feature, the vehicle interaction feature and the second lane segment fusion feature, an attention mechanism operation is performed to obtain a target fusion feature; based on the target fusion feature, a predicted trajectory of the vehicle to be predicted is generated.

[0009] Based on the above technical solution, the embodiment of the present application integrates the second lane segment fusion features and vehicle interaction features corresponding to the surrounding vehicles on the basis of the above lane segment sequence features, so as to accurately characterize the relative position and dynamic relationship between the predicted vehicle and the surrounding vehicles, and through the attention mechanism operation, the model can better understand the vehicle behavior patterns in complex traffic scenes, thereby generating a more reliable predicted trajectory.

[0010] In some embodiments, obtaining a first position sequence feature of a vehicle to be predicted includes: obtaining historical position information of the vehicle to be predicted; processing the historical position information based on a first multi-layer perceptron MLP to obtain a second position sequence feature of the vehicle to be predicted; fusing the second position sequence feature with a time embedding parameter to obtain a third position sequence feature; and performing an attention mechanism operation on the third position sequence feature to obtain a first position sequence feature.

[0011] Based on the above technical solution, the embodiment of the present application uses the first multi-layer perceptron MLP to process the historical position information, which can extract the deep-level features of the vehicle movement (second position sequence features), which help to more accurately predict the future position of the vehicle. At the same time, the introduction of time embedding parameters enables the model to consider the influence of time factors, further improving the accuracy of the prediction.

[0012] In some embodiments, obtaining lane segment sequence features of the vehicle to be predicted includes: obtaining surrounding lane segment position information of the vehicle to be predicted; and processing the surrounding lane segment position information based on a second multi-layer perceptron MLP to obtain lane segment sequence features.

[0013] Based on the above technical solution, the embodiment of the present application processes the surrounding lane segment position information through the second multi-layer perceptron MLP, and can efficiently extract the key features of the lane segment. As a feedforward neural network, the second multi-layer perceptron MLP has a strong nonlinear fitting ability and can learn the complex relationship between the lane segment position information and the vehicle movement, so that the model can more accurately understand the structure and constraints of the lane segment in subsequent position prediction or trajectory planning.

[0014] In some embodiments, an attention mechanism operation is performed based on the first position sequence features and the lane segment sequence features to obtain a first lane segment fusion feature, including: mapping the first position sequence features to a first query vector based on a third multi-layer perceptron MLP; mapping the first splicing feature to a first key vector and a first value vector based on a fourth multi-layer perceptron MLP; the first splicing feature is a feature obtained by splicing the first position sequence features and the lane segment sequence features; and performing an attention mechanism operation based on the first query vector, the first key vector, and the first value vector to obtain the first lane segment fusion feature.

[0015] Based on the above technical solution, the embodiment of the present application can flexibly map the first position sequence features and lane segment sequence features to query vectors, key vectors and value vectors in a high-dimensional space through the third multi-layer perceptron MLP and the fourth multi-layer perceptron MLP. In this way, the important information in the original features can be retained through this mapping method, while reducing redundancy and noise. At the same time, the attention mechanism can dynamically adjust the weights of different features in the fusion process according to the similarity between the query vector and the key vector, so that the trajectory prediction model can more accurately understand the relationship between the first position sequence features and the lane segment sequence features, thereby generating the first lane segment fusion feature.

[0016] In some embodiments, obtaining vehicle interaction features includes: obtaining surrounding vehicle position information of the vehicle to be predicted; and processing the surrounding vehicle position information based on a fifth multi-layer perceptron MLP to obtain vehicle interaction features.

[0017] Based on the above technical solution, the embodiment of the present application can establish a connection relationship between the predicted vehicle and the surrounding vehicles at the current moment, and use embedded representation and relative coordinates to extract edge features to achieve the construction of global interaction coding. This method can accurately capture the dynamic interaction relationship between each vehicle, and provides effective technical support for intelligent processing in the fields of intelligent transportation, autonomous driving, etc.

[0018] In some embodiments, an attention mechanism operation is performed based on the first lane segment fusion feature, the vehicle interaction feature, and the second lane segment fusion feature to obtain a target fusion feature, including: mapping the first lane interaction feature to a second query vector based on a sixth multi-layer perceptron MLP; mapping the second splicing feature to a second key vector and a second value vector based on a seventh multi-layer perceptron MLP; the second splicing feature is a feature obtained by splicing the first lane segment fusion feature, the vehicle interaction feature, and the second lane segment fusion feature; and performing an attention mechanism operation based on the second query vector, the second key vector, and the second value vector to obtain the target fusion feature.

[0019] According to a second aspect provided by the present application, a method for training a trajectory prediction model is provided, the method comprising: obtaining a first position sequence feature and a lane segment sequence feature of a vehicle at a first sample moment; the first position sequence feature characterizes the position of the vehicle at multiple historical moments before the first sample moment; the lane segment sequence feature characterizes the positional relationship between the vehicle and the surrounding lane segments at multiple historical moments; obtaining a target lane segment true value label and a true trajectory corresponding to the vehicle at the second sample moment; processing the first position sequence feature and the lane segment sequence feature based on a trajectory prediction model to obtain a predicted lane segment label and a predicted trajectory corresponding to the second sample moment; the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence feature and the lane segment feature to obtain a first lane segment fusion feature; based on the first lane segment fusion feature, obtain a predicted trajectory, and determine a predicted lane segment label corresponding to the predicted trajectory; calculate a first loss based on the target lane segment true value label and the predicted lane segment label; calculate a second loss based on the true trajectory and the predicted trajectory; and adjust the trajectory prediction model according to the first loss and the second loss.

[0020] Based on the above technical solution, the embodiment of the present application comprehensively considers the historical position information of the vehicle and the position relationship with the surrounding lane segments. The above two features work together in trajectory prediction, which can more comprehensively reflect the driving status and possible driving path of the vehicle. At the same time, through the operation of the attention mechanism, the model can dynamically adjust the weights of different features in the prediction process, so as to more accurately capture the key information of vehicle driving and improve the accuracy of prediction.

[0021] In some embodiments, obtaining a first position sequence feature and a lane segment sequence feature of a vehicle at a first sample moment includes: obtaining historical position information of the vehicle at multiple historical moments and surrounding lane segment position information of the to-be-predicted vehicle at multiple historical moments; processing the historical position information and the surrounding lane segment position information based on a trajectory prediction model to obtain a first position sequence feature and a lane segment sequence feature, respectively.

[0022] In some embodiments, the method also includes: obtaining vehicle interaction features of the vehicle and second lane segment fusion features corresponding to the surrounding vehicles of the vehicle; the vehicle interaction features are used to characterize the positional relationship between the vehicle and the surrounding vehicles at the first sample moment; the first position sequence features and the lane segment sequence features are processed based on the trajectory prediction model to obtain the predicted lane segment label and predicted trajectory corresponding to the second sample moment, including: processing the first lane segment fusion features, the vehicle interaction features and the second lane segment fusion features based on the trajectory prediction model to obtain the predicted lane segment label and predicted trajectory corresponding to the second sample moment; the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence features and the lane segment features to obtain the first lane segment fusion features; the attention mechanism operation is performed based on the first lane segment fusion features, the vehicle interaction features and the second lane segment fusion features to obtain the target fusion features; based on the target fusion features, the predicted trajectory of the vehicle is generated, and the predicted lane segment label corresponding to the predicted trajectory is determined.

[0023] According to a third aspect provided by the present application, a vehicle is provided, comprising a trajectory prediction device.

[0024] According to the fourth aspect provided by the present application, an electronic device is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to execute instructions to implement the method of the above-mentioned first aspect and any possible implementation manner thereof.

[0025] According to the fifth aspect provided by the present application, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method in the above-mentioned first aspect and any possible implementation method thereof.

[0026] According to the sixth aspect provided by the present application, a computer program product is provided, the computer program product comprising computer instructions, and when the computer instructions are executed on an electronic device, the electronic device executes the method of the above-mentioned first aspect and any possible implementation manner thereof.

[0027] According to a seventh aspect provided by the present application, a trajectory prediction device is provided, the device comprising: a processing unit and an acquisition unit.

[0028] Therefore, the above technical features of the present application have the following beneficial effects:

[0029] (1) The embodiment of the present application integrates multi-dimensional information (first position sequence features and lane segment sequence features) and uses the attention mechanism operation to enable the model to automatically learn and emphasize the position sequence features and lane segment sequence features that have an important impact on the prediction results, thereby improving the accuracy and robustness of the prediction.

[0030] (2) The embodiment of the present application integrates the second lane segment fusion features and vehicle interaction features corresponding to the surrounding vehicles on the basis of the above-mentioned lane segment sequence features, so as to accurately characterize the relative position and dynamic relationship between the predicted vehicle and the surrounding vehicles. Through the operation of the attention mechanism, the model can better understand the vehicle behavior patterns in complex traffic scenarios, thereby generating a more reliable predicted trajectory.

[0031] (3) The embodiment of the present application uses the first multi-layer perceptron MLP to process the historical position information, which can extract the deep-level features of the vehicle movement (second position sequence features), which help to more accurately predict the future position of the vehicle. At the same time, the introduction of time embedding parameters enables the model to consider the influence of time factors, further improving the accuracy of the prediction.

[0032] (4) The embodiment of the present application processes the surrounding lane segment position information through the second multi-layer perceptron MLP, and can efficiently extract the key features of the lane segment. The second multi-layer perceptron MLP, as a feedforward neural network, has a strong nonlinear fitting ability and can learn the complex relationship between the lane segment position information and the vehicle motion, so that the model can more accurately understand the structure and constraints of the lane segment in subsequent position prediction or trajectory planning.

[0033] (5) The embodiment of the present application can flexibly map the first position sequence features and the lane segment sequence features to the query vector, key vector and value vector in the high-dimensional space through the third multi-layer perceptron MLP and the fourth multi-layer perceptron MLP. In this way, the important information in the original features can be retained through this mapping method, while reducing redundancy and noise. At the same time, the attention mechanism can dynamically adjust the weights of different features in the fusion process according to the similarity between the query vector and the key vector, so that the trajectory prediction model can more accurately understand the relationship between the first position sequence features and the lane segment sequence features, thereby generating the first lane segment fusion feature.

[0034] (6) The embodiment of the present application realizes the construction of global interaction coding by establishing the connection relationship between the predicted vehicle and the surrounding vehicles at the current moment, and extracting edge features using embedded representation and relative coordinates. This method can accurately capture the dynamic interaction relationship between each vehicle and provide effective technical support for intelligent processing in the fields of intelligent transportation, autonomous driving, etc.

[0035] (7) The embodiment of the present application comprehensively considers the historical position information of the vehicle and the position relationship with the surrounding lane segments. The above two features work together in trajectory prediction, which can more comprehensively reflect the driving status and possible driving path of the vehicle. At the same time, through the operation of the attention mechanism, the model can dynamically adjust the weights of different features in the prediction process, thereby more accurately capturing the key information of vehicle driving and improving the accuracy of prediction.

[0036] It should be noted that the technical effects brought about by any implementation method in the second to seventh aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here.

[0037] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application, and do not constitute an improper limitation on the present application.

[0039] Figure 1 is a schematic structural diagram of a vehicle according to an exemplary embodiment;

[0040] Figure 2 is a block diagram of a trajectory prediction system according to an exemplary embodiment;

[0041] Figure 3 is a flow chart of a trajectory prediction method according to an exemplary embodiment;

[0042] Figure 4 is a scene diagram of a trajectory prediction method according to an exemplary embodiment;

[0043] Figure 5 is a scene diagram of another trajectory prediction method according to an exemplary embodiment;

[0044] Figure 6 is a scene diagram of another trajectory prediction method according to an exemplary embodiment;

[0045] Figure 7 is a flowchart of a method for training a trajectory prediction model according to an exemplary embodiment;

[0046] Figure 8 is a scene diagram of another trajectory prediction method according to an exemplary embodiment;

[0047] Fig. 9is a block diagram of another trajectory prediction device according to an exemplary embodiment;

[0048] Fig.10 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0049] In order to enable ordinary persons in the art to better understand the technical solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.

[0051] In the embodiments of the present application, words such as "exemplary", "for example", or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary", "for example", or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example", or "for example" is intended to present related concepts in a specific way.

[0052] First, the related technologies involved in this application are explained to facilitate understanding by those skilled in the art.

[0053] In recent years, the widespread application of deep neural networks in trajectory prediction has indeed brought significant progress to autonomous driving, intelligent transportation and other fields. Among them, the target-driven trajectory prediction method achieves a more accurate prediction of the vehicle's future trajectory by decomposing the problem into two subtasks: predicting possible targets and estimating motion based on contextual features. However, while pursuing low prediction errors, these methods often ignore the physical feasibility of the predicted trajectory, which is an issue worthy of in-depth exploration.

[0054] To solve the above problems, the goal can be to use a more complex trajectory representation method to more comprehensively describe the vehicle's occupancy in space, and then generate a predicted trajectory for the vehicle to overcome the problem of low feasibility. This method can more accurately reflect the actual space occupied by the vehicle, thereby improving the physical feasibility of the prediction results.

[0055] In summary, although deep learning-based trajectory prediction methods have achieved remarkable results in reducing prediction errors, physical feasibility is still one of the key factors restricting its practical application. In the future, researchers need to introduce more physical constraints and more complex trajectory representation methods into the prediction model to ensure that the generated prediction trajectory is both accurate and feasible. This will provide more reliable guarantees for the practical application of intelligent transportation systems such as autonomous driving.

[0056] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.

[0057] The trajectory prediction method provided in the embodiment of the present application can be applied in a vehicle. Figure 1 , which is a schematic diagram of a vehicle structure provided by an embodiment of the present application, the vehicle 100 may include a chassis 110, a body 120, and wheels 130. It is understandable that the vehicle 100 may be a fuel vehicle, an electric vehicle, a hybrid vehicle, a gas vehicle, a methanol vehicle, a solar vehicle, etc.

[0058] For example, the vehicle 100 may be a passenger vehicle such as a sedan, a sport utility vehicle (SUV), a multi-purpose vehicle (MPV), or a bus, a truck, a semi-trailer, etc. This application does not impose any specific restrictions on this.

[0059] It is understandable that the above components are merely examples of some components of the vehicle 100 and are not limitations on the specific structure of the vehicle 100 .

[0060] Optionally, in order to control the vehicle, the vehicle 100 may further include a trajectory prediction system 140. The trajectory prediction system 140 may predict the motion trajectory of the vehicle to be predicted around the vehicle 100.

[0061] like Figure 2 As shown, the present embodiment provides a block diagram of a trajectory prediction system 140. The trajectory prediction system 140 includes an acquisition module 210 and a trajectory prediction module 220.

[0062] The acquisition module 210 is used to acquire the first position sequence features and lane segment sequence features of the vehicle to be predicted; and send the first position sequence features and lane segment sequence features to the trajectory prediction module 220.

[0063] In some embodiments, the first position sequence feature is used to characterize the position features of the vehicle to be predicted at multiple historical moments; the lane segment sequence feature is used to characterize the positional relationship between the vehicle to be predicted and the surrounding lane segments.

[0064] The trajectory prediction module 220 is used to perform an attention mechanism operation based on the first position sequence feature and the lane segment sequence feature to obtain the first lane segment fusion feature, and obtain the predicted trajectory of the vehicle to be predicted based on the first lane segment fusion feature.

[0065] It should be noted that the trajectory prediction system described in the embodiment of the present application is to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. A person of ordinary skill in the art will know that with the evolution of electronic devices and the emergence of other electronic devices, the technical solution provided by the embodiment of the present application is also applicable to similar technical problems. The methods in the following embodiments can all be implemented in a trajectory prediction system having the above hardware structure.

[0066] The methods in the following embodiments can all be implemented in a trajectory prediction system having the above hardware structure.

[0067] The trajectory prediction method provided by the embodiment of the present application is described in detail below with reference to the accompanying drawings.

[0068] The trajectory prediction method of the embodiment of the present application can be applied to predict the motion trajectory of a vehicle to be predicted. Figure 3 As shown, the trajectory prediction method may include steps 301 to 303. Step 301 may also be referred to as a process of "obtaining the first position sequence characteristics and lane segment sequence characteristics of the vehicle to be predicted", and steps 302 to 303 may be referred to as a process of "generating a predicted trajectory of the vehicle to be predicted". Steps 301 to 303 are described in detail below.

[0069] Step 301: Obtain first position sequence features and lane segment sequence features of a vehicle to be predicted.

[0070] In an embodiment related to the present application, the first position sequence feature is used to characterize the position features of the vehicle to be predicted at multiple historical moments; the lane segment sequence feature is used to characterize the positional relationship between the vehicle to be predicted and the surrounding lane segments.

[0071] In some embodiments, the trajectory prediction system can obtain the historical position information of the vehicle to be predicted, and process the historical position information based on the first multi-layer perceptron MLP to obtain the second position sequence characteristics of the vehicle to be predicted, and then fuse the second position sequence characteristics and the time embedding parameters to obtain the third position sequence characteristics, perform attention mechanism operations on the third position sequence characteristics, and obtain the first position sequence characteristics.

[0072] For example, Figure 4 As shown, the trajectory prediction system can obtain the historical position information of the vehicle to be predicted (H i,t ), each historical location information includes but is not limited to a timestamp and the coordinate value of the vehicle to be predicted, where t represents different historical timestamps (t=0, t=-1, t=-2), and i represents the i-th vehicle to be predicted.

[0073] Furthermore, the trajectory prediction system can check the historical position information H of the vehicle to be predicted. i,t Whether the corresponding timestamp is valid. For location information with invalid timestamps (such as missing, incorrect format, etc.) or obviously abnormal compared with other historical location information (such as too large a time jump), it is regarded as an invalid location and its corresponding coordinate value is set to zero to avoid its influence in subsequent analysis.

[0074] For valid position data, calculate the coordinate increment ΔH between adjacent position information (eg, step size is 1) i,t , that is, Δx=x(t+1)-x(t) and Δy=y(t+1)-y(t) (or Δz in three-dimensional space), which represents the instantaneous direction and instantaneous speed information of the vehicle to be predicted.

[0075] In addition, the instantaneous speed information can also be combined with the timestamp information to calculate the moving distance of the predicted vehicle between adjacent positions and divide it by the time interval to obtain the instantaneous speed information. The calculated instantaneous direction and instantaneous speed information can be output for subsequent path planning, obstacle avoidance, navigation and other applications.

[0076] Therefore, the embodiment of the present application effectively improves the accuracy of the instantaneous direction and instantaneous speed information of the vehicle to be predicted through the timestamp validity check and invalid position data processing mechanism. At the same time, the coordinate increment with a step size of 1 is used to represent the motion state of the vehicle to be predicted, providing reliable data support for the motion analysis and control of the vehicle to be predicted.

[0077] Alternatively, if Figure 5 As shown, the embodiment of the present application can construct a temporal encoder through a Transformer model, and calculate the self-attention feature and the fully connected feature respectively by constructing two residual modules.

[0078] Furthermore, the trajectory prediction system can project the coordinate increment ΔHi,t into a feature space of dimension d as the second position sequence feature e through a first multilayer perceptron (MLP). i,t .

[0079] Furthermore, in the embodiment of the present application, a set of learnable parameters can be initialized as a time embedding, whose length is equal to the time step of the historical position information of the vehicle to be predicted, and the initialized time embedding is spliced ​​with the second position feature sequence. The splicing operation can be performed on the feature dimension, that is, the time embedding is used as an additional dimension of the feature of the vehicle to be predicted.

[0080] Then, the concatenated features are transformed in dimension to meet the needs of subsequent prediction trajectory model processing. After time embedding initialization, feature concatenation and dimension transformation, the third position sequence features containing time series information are obtained. The dimension transformation may include adjusting the shape of the features, performing linear transformation, etc.

[0081] It should be understood that the dimension of the time embedding can be set according to the specific task requirements, and the parameters of the time embedding will be continuously optimized in the subsequent model training process to better capture the timing information.

[0082] Furthermore, the third position sequence feature is mapped into query vector Query, key vector Key and value vector Value features through layer normalization and the first multi-layer perceptron, that is, the following formula 1:

[0083] {Q t ,K t ,V t}={φ Q (e i,t ),φ K (e i,t ),φ V (e i,t )} Formula 1

[0084] Among them, φ Q (e i,t ) represents the query vector Query feature mapped to the third position sequence feature, φ K (e i,t ) represents the key vector Key feature of the third position sequence feature mapping, φ V (e i,t ) represents the value vector Value feature mapped to the third position sequence feature.

[0085] The first position sequence feature is calculated using the self-attention mechanism, which is the following formula 2:

[0086]

[0087] in, represents the first position sequence feature, i represents the i-th vehicle to be predicted, t1 and t2 represent different historical moments, represents the transpose of the query vector Query feature matrix of the vehicle to be predicted at time t1, T represents the transpose of the matrix, Represents the key vector Key feature matrix of the vehicle to be predicted at time t2.

[0088] Preferably, if Figure 5 As shown in Figure 1, the trajectory prediction system can use multi-head attention for stable training, that is, Q (query vector Query feature), K (key vector Key feature) and V (value vector Value feature) are divided into g groups, so that the batch size is expanded to g times, and the dimension of each group of features is After completing the attention calculation, the final attention feature is obtained by concatenating multiple features.

[0089] Furthermore, the attention features are transformed in turn through layer normalization and multi-layer perceptron to obtain the fully connected features, that is, the following formula 3:

[0090]

[0091] in, represents the fully connected feature, Represents the first position sequence feature.

[0092] Furthermore, a residual connection is introduced to obtain new features of the vehicle to be predicted, that is, the following formula 4:

[0093]

[0094] in, represents the first position sequence feature, represents the position sequence feature obtained by the self-attention mechanism, e i,t represents the second position sequence feature, is a fully connected feature.

[0095] In order to reduce the overfitting effect, random depth and layer scaling are introduced in the residual connection. The random depth means that the residual connection has a certain probability of being abandoned during the training process of the model. This probability increases with the increase of the number of layers, that is, the following formula 5:

[0096]

[0097] In the formula, p l is the abandonment probability of the lth layer, L is the total number of layers, p L is the abandonment probability of the last layer. Layer scaling is to independently scale the residual features of each layer in the feature dimension, so that the differences in features of different channels are richer, that is, the following formula 6:

[0098] e′=w ls·e Formula 6

[0099] In the formula, e and e′ are the original residual features and the updated residual features respectively, w ls is the learnable weight of .

[0100] It should be understood that Figure 4 As shown, the first position sequence feature of the vehicle to be predicted is output by the time series encoder. The time series encoder can encode the historical state or behavior data of the vehicle to be predicted and extract features with time series characteristics, which represent the state or behavior pattern of the vehicle to be predicted at the current moment.

[0101] In some embodiments, the trajectory prediction system may obtain the surrounding lane segment position information of the vehicle to be predicted, and process the surrounding lane segment position information based on a second multi-layer perceptron MLP to obtain lane segment sequence features.

[0102] Exemplarily, the trajectory prediction system can determine the lane segment (i.e., the surrounding lane segment) at a preset distance in the direction of movement of the vehicle to be predicted based on the historical position information of the vehicle to be predicted, and process the coordinates of the screened surrounding lane segment according to the second multi-layer perceptron MLP to obtain the position information of the surrounding lane segment.

[0103] Furthermore, the position information of the surrounding lane segments (i.e., the midpoint coordinates of the line segments of the surrounding lane segments) is converted to the local coordinate system of the vehicle to be predicted, so that the position of the midpoint coordinates of the line segments relative to the vehicle to be predicted can be clearly expressed.

[0104] The second multi-layer perceptron is used to process the transformed line segment midpoint coordinates. The second multi-layer perceptron extracts deep features in the line segment midpoint coordinates through nonlinear transformation of the input data, and the deep features are used as the features of the edges in the graph structure (i.e., lane segment sequence features).

[0105] Preferably, the embodiment of the present application can combine the lane segment sequence features processed by the second multi-layer perceptron MLP with the first position sequence features of the vehicle to be predicted to construct a spatiotemporal graph structure between the lane segment and the vehicle to be predicted. The spatiotemporal graph structure not only reflects the spatial relationship between the lane segment and the vehicle to be predicted, but also incorporates the temporal state information of the vehicle to be predicted, thereby achieving a comprehensive description of the surrounding environment of the vehicle to be predicted.

[0106] The embodiment of the present application uses the above method to efficiently construct the relative position relationship (space-time graph structure) between the lane segment and the vehicle to be predicted, providing strong support for subsequent path planning, trajectory prediction and other tasks.

[0107] Step 302: Perform an attention mechanism operation based on the first position sequence features and the lane segment sequence features to obtain the first lane segment fusion features.

[0108] In some embodiments, the trajectory prediction system can map the first position sequence feature to a first query vector based on a third multi-layer perceptron MLP, and map the first concatenated feature to a first key vector and a first value vector based on a fourth multi-layer perceptron MLP, and then perform an attention mechanism operation based on the first query vector, the first key vector, and the first value vector to obtain a first lane segment fusion feature.

[0109] The first splicing feature is a feature obtained by splicing the first position sequence feature and the lane segment sequence feature;

[0110] For example, Figure 6 As shown, the trajectory prediction system can project the first position sequence feature into Q (the first query vector Q) through the third multi-layer perceptron MLP i,t ), the first concatenated feature after the first position sequence feature and the lane segment feature (lane segment sequence) are concatenated by the fourth multi-layer perceptron MLP is projected into K (the first key vector K j,t ) and V(the first value vector V j,t ). Wherein, the number of channels can be d.

[0111] like Figure 4 As shown in the figure, the lane segment sequence features are fused into the first position sequence features using the lane feature fusion module’s heterogeneous attention mechanism to obtain the residual features. That is, the following formula 7:

[0112]

[0113] Furthermore, the trajectory prediction system can use the residual features After being concatenated with the first position sequence feature, it is normalized and processed by the fourth multi-layer perceptron MLP to obtain the updated position sequence feature e of the vehicle to be predicted. i ′, that is, the following formula 8:

[0114]

[0115] Among them, η1 and η2 represent normalization functions, and φ represents the fourth multi-layer perceptron MLP network structure of the fully connected layer.

[0116] The directional lane segment features are further correlated with the vehicle A to be predicted through the temporal encoder. i The node features of the vehicle sequence to be predicted at t = 0 are fused and the i,t |t=0 is embedded as the vehicle to be predicted b i .

[0117] Furthermore, the lane segment sequence features and the updated position sequence features of the vehicle to be predicted are transformed into each other using a temporal encoder.i ′ is fused. On the basis of time series coding, the feature of the vehicle sequence to be predicted at the specific moment of time t=0 is selected. This feature contains both the dynamic information of the vehicle to be predicted itself and the static and dynamic features of the lane segment sequence characteristics. This feature is set as the embedding representation b of the vehicle to be predicted i , the embedding vector b i It can comprehensively and compactly reflect the state of the vehicle to be predicted at the initial moment and its interaction with the surrounding environment.

[0118] It should be understood that the time series encoder can capture the dynamic changes of time series data and effectively integrate the information in the time dimension through its internal mechanism (such as recurrent neural network RNN, long short-term memory network LSTM or gated recurrent unit GRU, etc.), thereby generating a new feature set that integrates spatiotemporal characteristics.

[0119] Step 303: Obtain a predicted trajectory of the vehicle to be predicted based on the fusion features of the first lane segment.

[0120] In some embodiments, the trajectory prediction system can obtain vehicle interaction features and second lane segment fusion features corresponding to the surrounding vehicles of the vehicle to be predicted, and then perform attention mechanism operations based on the first lane segment fusion features, the vehicle interaction features and the second lane segment fusion features to obtain target fusion features, and generate a predicted trajectory of the vehicle to be predicted based on the target fusion features.

[0121] Among them, the vehicle interaction feature is used to characterize the positional relationship between the vehicle to be predicted and the surrounding vehicles.

[0122] For example, Figure 4 As shown, the trajectory prediction module of the trajectory prediction system can input the target fusion features into multiple groups of multi-layer perceptrons for decoding, obtain multiple predicted trajectories and the confidence corresponding to each predicted trajectory, and determine the predicted trajectory with the highest confidence among the multiple predicted trajectories as the predicted trajectory of the vehicle to be predicted.

[0123] In some embodiments, the target fusion features mentioned above can be obtained in the following manner: the trajectory prediction system can map the first lane interaction feature to a second query vector based on the sixth multi-layer perceptron MLP, and map the second splicing feature to a second key vector and a second value vector based on the seventh multi-layer perceptron MLP, and then perform an attention mechanism operation based on the second query vector, the second key vector and the second value vector to obtain the target fusion features.

[0124] The second splicing feature is a feature obtained by splicing the first lane segment fusion feature, the vehicle interaction feature and the second lane segment fusion feature.

[0125] For example, Figure 4As shown, the lane target prediction module of the trajectory prediction system can project the first lane interaction feature into the second query vector Q through the sixth multi-layer perceptron MLP i,t , the second concatenated feature after the first lane segment fusion feature, the vehicle interaction feature and the second lane segment fusion feature are concatenated through the seventh multi-layer perceptron MLP is projected as the second key vector K l,t and the second value vector V l,t . Wherein, l represents the lth surrounding vehicle (l=1, 2, ...), and the number of channels can be d.

[0126] The self-attention mechanism is used to fuse the second lane segment fusion feature into the first position sequence feature to obtain the residual feature That is, the following formula 9:

[0127]

[0128] Furthermore, the trajectory prediction system can use the residual features After being concatenated with the first position sequence feature, it is normalized and processed by the seventh multi-layer perceptron MLP to obtain the target fusion feature b i ′, that is, the following formula 10:

[0129]

[0130] Among them, η1 and η2 represent normalization functions, and φ represents the seventh multi-layer perceptron MLP network structure of the fully connected layer.

[0131] In some embodiments, the above-mentioned vehicle interaction features can be obtained in the following manner: the trajectory prediction system can obtain the surrounding vehicle position information of the vehicle to be predicted, and process the surrounding vehicle position information based on the fifth multi-layer perceptron MLP to obtain the vehicle interaction features.

[0132] For example, Figure 4 As shown, the interactive fusion module of the trajectory prediction system can obtain the position information of each vehicle around the vehicle to be predicted, and the position information includes but is not limited to the historical behavior, attributes and current status of the vehicle, and then calculate the relative coordinates of the surrounding vehicle position information relative to the vehicle to be predicted. The relative coordinates can reflect the spatial position relationship between the vehicle to be predicted and the surrounding vehicles. The relative coordinates are processed by the fifth multi-layer perceptron MLP to obtain the vehicle interaction characteristics of the vehicle to be predicted and the surrounding vehicles. The vehicle interaction characteristics not only include the spatial distance information between the vehicle to be predicted and the surrounding vehicles, but also can capture more complex interaction relationships, such as speed differences and movement directions, through the nonlinear mapping capability of the fifth multi-layer perceptron.

[0133] Optionally, the trajectory prediction system can also construct a global interaction code through the vehicle interaction features obtained above, the location information of the vehicle to be predicted and the surrounding vehicles. That is to say, the location information of the vehicle to be predicted, the surrounding vehicles and the vehicle interaction features are integrated into a graph structure, in which the vehicle to be predicted is used as a node and the connection relationship between them is used as an edge. Through graph neural networks (GNN) or other graph processing algorithms, deep information in the global interaction code can be further extracted to provide support for subsequent decision-making, prediction or optimization tasks.

[0134] It should be understood that the embodiment of the present application realizes the construction of global interaction coding by establishing the connection relationship between the predicted vehicle and the surrounding vehicles at the current moment, and extracting edge features using embedded representation and relative coordinates. This method can accurately capture the dynamic interaction relationship between each vehicle and provide effective technical support for intelligent processing in the fields of intelligent transportation, autonomous driving, etc.

[0135] The trajectory prediction method provided in the embodiment of the present application can capture the dynamic driving mode of the vehicle by comprehensively considering the vehicle position characteristics (i.e., the first position sequence characteristics) at multiple historical moments, thereby more accurately predicting its future trajectory. The lane segment sequence characteristics are introduced, i.e., the positional relationship between the vehicle and the surrounding lane segments, which provides additional contextual information, helps to more accurately understand the vehicle's driving environment and further improve the accuracy of the prediction.

[0136] At the same time, the trajectory prediction method provided in the embodiment of the present application relies on the dynamic driving mode of the vehicle and its relative position to the lane segment, rather than a fixed path or mode, which enables the model to adapt to different traffic scenarios and road structures, and the introduction of the attention mechanism enables the model to dynamically focus on important input features, thereby making more reasonable predictions in different situations.

[0137] like Figure 7 As shown, an embodiment of the present application provides a method for training a trajectory prediction model, which specifically includes the following steps 701-706.

[0138] Step 701: Obtain a first position sequence feature and a lane segment sequence feature of a vehicle at a first sample time.

[0139] In an embodiment related to the present application, the first position sequence feature represents the position of the vehicle at multiple historical moments before the first sample moment; the lane segment sequence feature represents the position relationship between the vehicle and the surrounding lane segments at multiple historical moments.

[0140] In some embodiments, historical position information of the vehicle at multiple historical moments and surrounding lane segment position information of the vehicle to be predicted at multiple historical moments are obtained, and the historical position information and the surrounding lane segment position information are processed based on a trajectory prediction model to obtain a first position sequence feature and a lane segment sequence feature, respectively.

[0141] It should be noted that the process of acquiring the trajectory prediction model and processing the historical position information and the surrounding lane segment position information is the same as that of the above-mentioned trajectory prediction method, and will not be repeated here.

[0142] Step 702: Obtain the target lane segment true value label and true trajectory corresponding to the vehicle at the second sample time.

[0143] In an embodiment of the present application, a given trajectory prediction data set can be divided into a training set and a validation set in a ratio of 4:1. The training set accounts for 80% of the total data volume and is used to train the trajectory prediction model; the validation set accounts for 20% of the total data volume and is used to evaluate the prediction performance of the model. The true trajectory corresponding to the second sample moment of the vehicle included in the data set is set as the final prediction output true value label of the trajectory prediction model. It should be understood that the true value label can be used as supervisory information to guide and optimize the prediction task training process of the model.

[0144] Exemplarily, lane data in the trajectory prediction data set is preprocessed. The preprocessing step is to filter according to the current driving direction of the vehicle, that is, to search for surrounding lane segments within a certain distance range along the driving direction of the vehicle, and the surrounding lane segments can be regarded as future candidate target vectors.

[0145] Within a certain distance range, the surrounding lane segments that match the vehicle's driving direction are screened out to form a set of future candidate target vectors. Then, in the candidate target vector set, the distance between the vehicle's trajectory point and each surrounding lane segment at each moment is calculated, and the closest target surrounding lane segment is selected as the target lane segment label at that moment. The target lane segment label constitutes the true value label set of the target lane segment, which is used to supervise the training of the auxiliary task detection head.

[0146] Step 703: Process the first position sequence features and the lane segment sequence features based on the trajectory prediction model to obtain a predicted lane segment label and a predicted trajectory corresponding to the second sample time.

[0147] In an embodiment related to the present application, the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence features and the lane segment features to obtain the first lane segment fusion features; based on the first lane segment fusion features, a predicted trajectory is obtained, and a predicted lane segment label corresponding to the predicted trajectory is determined.

[0148] It should be noted that the process of determining the first lane segment fusion feature and obtaining the predicted trajectory according to the first lane segment fusion feature in the above trajectory prediction model is the same as the process of the trajectory prediction method, which will not be repeated here. It should be understood that after obtaining the predicted trajectory, the predicted trajectory traveled by the vehicle will correspond to the corresponding lane segment, and the trajectory prediction model can use the corresponding lane segment as the predicted lane segment label.

[0149] Preferably, if Figure 8 As shown, the trajectory prediction model can map the first position sequence feature into a third query vector according to the multi-layer perceptron MLP, and map the lane segment feature into a third key vector and a third value vector according to the multi-layer perceptron MLP, that is, the following formula 11:

[0150] {Q i,t}={(φ Q (e i,t )}

[0151] {K j,t ,V j,t}={(φ K (e j,t ),φ V (e j,t )} Formula 11

[0152] In the formula, φ Q (e i,t ) represents the third query vector Query feature mapped to the first position sequence feature, φ K (e j,t ) represents the third key vector Key feature of the lane segment feature map, φ V (e j,t ) represents the third value vector Value feature of the lane segment feature mapping.

[0153] The third key vector, the third value vector, and the third query vector are processed by the multi-head cross attention mechanism to determine the lane segment with the highest classification score in the lane segment set as the predicted lane segment. In other words, the multi-head cross attention mechanism is used to calculate the multi-head attention coefficient of the vehicle to the surrounding lane segments at each historical moment, that is, the following formula 12:

[0154]

[0155] Among them, α i,j,t represents the cross attention coefficient, cross_attn represents the single-head cross attention calculation function, represents the transpose of the first position sequence feature of the i-th vehicle at time t, K j,t and V j,t They represent the lane segment sequence features at time t respectively.

[0156] Furthermore, for each head, the corresponding cross attention coefficient is summed, and then the summation result is normalized by applying the Softmax function. It should be understood that this step is used to calculate the classification score μ_(j, t_f) for each lane segment at the second sample time according to the attention weight of each head.

[0157] Using the target lane segment label at each moment, the target lane segment classification score corresponding to the target lane segment label is selected from the calculated classification scores. This step ensures the consistency between the classification score used for subsequent loss calculation and the actual lane target segment.

[0158] In some embodiments, the trajectory prediction model may obtain vehicle interaction features of the vehicle and second lane segment fusion features corresponding to surrounding vehicles of the vehicle.

[0159] The vehicle interaction feature is used to characterize the positional relationship between the vehicle and surrounding vehicles at the first sample moment.

[0160] It should be noted that the above process of obtaining the vehicle interaction features and the second lane segment fusion features is the same as the process of obtaining the vehicle interaction features and the second lane segment fusion features in the trajectory prediction method, and will not be repeated here.

[0161] In some embodiments, the first lane segment fusion feature, the vehicle interaction feature, and the second lane segment fusion feature are processed based on the trajectory prediction model to obtain a predicted lane segment label and a predicted trajectory corresponding to the second sample time.

[0162] Among them, the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence features and the lane segment features to obtain the first lane segment fusion features; perform a hetero-attention mechanism operation based on the first lane segment fusion features, the vehicle interaction features and the second lane segment fusion features to obtain the target fusion features; based on the target fusion features, generate the predicted trajectory of the vehicle, and determine the predicted lane segment label corresponding to the predicted trajectory.

[0163] It should be noted that the process of determining the first lane segment fusion feature and the target fusion feature by the above trajectory prediction model, and obtaining the predicted trajectory according to the target fusion feature is the same as the process of the trajectory prediction method, which will not be repeated here. It should be understood that after obtaining the predicted trajectory, the predicted trajectory traveled by the vehicle will correspond to the corresponding lane segment, and the trajectory prediction model can use the corresponding lane segment as the predicted lane segment label.

[0164] Step 704: Calculate a first loss based on the target lane segment true value label and the predicted lane segment label.

[0165] In some embodiments, based on the classification score of the selected target lane segment, a cross entropy classification loss (BCE_loss) is used as a loss function to calculate the training loss of the auxiliary task. This loss function is used to measure the error between the classification score predicted by the model and the actual lane target point, thereby guiding the training process of the model.

[0166] Exemplarily, the target lane segment true value label is matched with the predicted lane segment label, for example, based on the time or space dimension, ensuring that the two are compared at the same time point or spatial position, and the difference between the matched target lane segment true value label and the predicted lane segment label is calculated, and the difference includes but is not limited to the position deviation, direction deviation, width deviation, etc. of the lane segment.

[0167] Furthermore, based on the above difference, the first loss function loss is calculated. goal , which is the cross entropy target classification loss (denoted as loss_goal). This loss function aims to quantify the inconsistency between the predicted lane segment label and the target lane segment true value label. It can usually adopt mathematical distance metrics (such as Euclidean distance, Manhattan distance, etc.) or probability metrics (such as cross entropy loss, KL divergence, etc.).

[0168] It should be understood that the calculation result of the first loss function reflects the accuracy of the model prediction, that is, the smaller the difference between the lane segment predicted by the model and the actual lane segment, the better the prediction performance of the model.

[0169] Step 705: Calculate a second loss based on the actual trajectory and the predicted trajectory.

[0170] Exemplarily, the actual trajectory and the predicted trajectory are synchronized in time or aligned in space to ensure that the two are compared at the same time point or spatial position. For example, an appropriate measurement method is used to calculate the difference between the actual trajectory and the predicted trajectory. The difference includes but is not limited to differences in the geometric shape of the trajectory (such as path length, curvature, etc.), dynamic characteristics (such as speed, acceleration changes, etc.), and time synchronization differences (such as time delay, time interval, etc.).

[0171] Based on the above differences, the second loss function is calculated. The second loss function aims to quantify the inconsistency between the predicted trajectory and the true trajectory, usually using mathematical distance metrics (such as Euclidean distance, Manhattan distance, Fréchet distance, etc.), dynamic characteristic metrics (such as velocity error, acceleration error, etc.) or time synchronization metrics (such as time delay loss, etc.).

[0172] In another exemplary embodiment, the predicted trajectory and the true trajectory are subjected to the SmoothL1 regression loss loss_reg, and the predicted trajectory is subjected to the second loss function loss_cls (cross entropy classification loss). The cross entropy classification loss is intended to improve the confidence level corresponding to the modality so that it more accurately reflects the true situation.

[0173] Step 706: Adjust the trajectory prediction model according to the first loss and the second loss.

[0174] Exemplarily, the calculated first loss and second loss are fused to form a joint loss function. For example, this can be achieved by a simple weighted summation, product or other complex nonlinear combination, wherein the joint loss function aims to fully reflect the performance of the model in both lane segment prediction and trajectory prediction.

[0175] Furthermore, the parameters of the trajectory prediction model are adjusted based on the gradient information of the joint loss function using gradient descent or other optimization algorithms. During the adjustment process, the model will continuously try to reduce the joint loss, thereby gradually improving the accuracy of lane segment prediction and trajectory prediction.

[0176] It should be understood that during the model training phase, the adjusted model can be iteratively trained on a dataset containing a rich variety of scenarios while continuously monitoring the changes in the joint loss. Through continuous parameter adjustment and optimization, the model gradually converges to the optimal solution.

[0177] In the model validation phase, an independent validation data set is used to evaluate the performance of the trained model. By comparing the difference between the model's prediction results on the validation set and the actual data, the model's generalization ability and prediction accuracy are verified. Then, based on the validation results, the model is further optimized and adjusted.

[0178] In another exemplary embodiment, the final training loss function loss_total is the first loss function loss goal , a linear combination of the SmoothL1 regression loss loss_reg and the second loss function loss_cls. Specifically, loss_total can be calculated as Formula 13:

[0179] loss_total=(1-α)·[β·loss_cls+(1-β)·loss_reg]+α·loss_goal Formula 13

[0180] Among them, the linear coefficients α and β are set to 0.2 and 0.5 respectively to balance the influence of different loss terms and achieve the effect of lane target prediction assisted trajectory prediction. During the training process, the embodiment of the present application uses the back propagation technology in the neural network model training to iteratively update the model parameters. By continuously adjusting the model parameters, the calculation result of the loss function loss_total continues to decrease. When the calculation result of the loss function is less than a certain threshold, it is considered that the model has been trained to the optimal mode, and the model has a higher prediction accuracy and reliability.

[0181] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to achieve the above functions, the image acquisition device or electronic device includes a hardware structure and / or software module corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0182] The embodiment of the present application can exemplarily divide the functional modules of the image acquisition device or electronic device according to the above method. For example, the image acquisition device or electronic device can include various functional modules corresponding to the various functional divisions, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0183] Fig. 9 is a block diagram of another trajectory prediction device according to an exemplary embodiment. The trajectory prediction device can be applied to a trajectory prediction method or a training method. Fig. 9 , the trajectory prediction device includes: a processing unit 901 and an acquisition unit 902. The acquisition unit 902 is used to acquire the first position sequence feature and lane segment sequence feature of the vehicle to be predicted; the first position sequence feature is used to characterize the position feature of the vehicle to be predicted at multiple historical moments; the lane segment sequence feature is used to characterize the position relationship between the vehicle to be predicted and the surrounding lane segments; the processing unit 901 is used to perform an attention mechanism operation based on the first position sequence feature and the lane segment sequence feature to obtain the first lane segment fusion feature; the processing unit 901 is also used to obtain the predicted trajectory of the vehicle to be predicted based on the first lane segment fusion feature.

[0184] In some embodiments, the processing unit 901 is specifically used to obtain vehicle interaction features and second lane segment fusion features corresponding to the surrounding vehicles of the vehicle to be predicted; the vehicle interaction features are used to characterize the positional relationship between the vehicle to be predicted and the surrounding vehicles; based on the first lane segment fusion features, the vehicle interaction features and the second lane segment fusion features, an attention mechanism operation is performed to obtain a target fusion feature; based on the target fusion feature, a predicted trajectory of the vehicle to be predicted is generated.

[0185] In some embodiments, the processing unit 901 is specifically used to obtain historical position information of the vehicle to be predicted; process the historical position information based on the first multi-layer perceptron MLP to obtain a second position sequence feature of the vehicle to be predicted; fuse the second position sequence feature and the time embedding parameter to obtain a third position sequence feature; perform an attention mechanism operation on the third position sequence feature to obtain a first position sequence feature.

[0186] In some embodiments, the processing unit 901 is specifically used to obtain the surrounding lane segment position information of the vehicle to be predicted; and process the surrounding lane segment position information based on the second multi-layer perceptron MLP to obtain the lane segment sequence features.

[0187] In some embodiments, the processing unit 901 is specifically used to map the first position sequence feature into a first query vector based on a third multi-layer perceptron MLP; map the first splicing feature into a first key vector and a first value vector based on a fourth multi-layer perceptron MLP; the first splicing feature is a feature obtained by splicing the first position sequence feature and the lane segment sequence feature; perform an attention mechanism operation based on the first query vector, the first key vector and the first value vector to obtain a first lane segment fusion feature.

[0188] In some embodiments, the processing unit 901 is specifically used to obtain the surrounding vehicle position information of the vehicle to be predicted; and process the surrounding vehicle position information based on the fifth multi-layer perceptron MLP to obtain vehicle interaction features.

[0189] In some embodiments, the processing unit 901 is specifically used to map the first lane interaction feature into a second query vector based on a sixth multi-layer perceptron MLP; map the second splicing feature into a second key vector and a second value vector based on a seventh multi-layer perceptron MLP; the second splicing feature is a feature obtained by splicing the first lane segment fusion feature, the vehicle interaction feature and the second lane segment fusion feature; perform an attention mechanism operation based on the second query vector, the second key vector and the second value vector to obtain a target fusion feature.

[0190] The acquisition unit 902 is also used to obtain the first position sequence features and lane segment sequence features of the vehicle at the first sample moment; the first position sequence features represent the position of the vehicle at multiple historical moments before the first sample moment; the lane segment sequence features represent the positional relationship between the vehicle and the surrounding lane segments at multiple historical moments; the acquisition unit 902 is also used to obtain the target lane segment true value label and the real trajectory corresponding to the vehicle at the second sample moment; the processing unit 901 is also used to process the first position sequence features and the lane segment sequence features based on the trajectory prediction model to obtain the predicted lane segment label and the predicted trajectory corresponding to the second sample moment; the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence features and the lane segment features to obtain the first lane segment fusion features; based on the first lane segment fusion features, a predicted trajectory is obtained, and a predicted lane segment label corresponding to the predicted trajectory is determined; a first loss is calculated based on the target lane segment true value label and the predicted lane segment label; a second loss is calculated based on the real trajectory and the predicted trajectory; and the trajectory prediction model is adjusted according to the first loss and the second loss.

[0191] In some embodiments, the processing unit 901 is specifically used to obtain historical position information of the vehicle at multiple historical moments and surrounding lane segment position information of the vehicle to be predicted at multiple historical moments; the historical position information and the surrounding lane segment position information are processed based on the trajectory prediction model to obtain the first position sequence feature and the lane segment sequence feature, respectively.

[0192] In some embodiments, the processing unit 901 is also used to obtain the vehicle interaction features of the vehicle and the second lane segment fusion features corresponding to the surrounding vehicles of the vehicle; the vehicle interaction features are used to characterize the positional relationship between the vehicle and the surrounding vehicles at the first sample moment; the first position sequence features and the lane segment sequence features are processed based on the trajectory prediction model to obtain the predicted lane segment label and predicted trajectory corresponding to the second sample moment, including: processing the first lane segment fusion features, the vehicle interaction features and the second lane segment fusion features based on the trajectory prediction model to obtain the predicted lane segment label and predicted trajectory corresponding to the second sample moment; the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence features and the lane segment features to obtain the first lane segment fusion features; the attention mechanism operation is performed based on the first lane segment fusion features, the vehicle interaction features and the second lane segment fusion features to obtain the target fusion features; according to the target fusion features, the predicted trajectory of the vehicle is generated, and the predicted lane segment label corresponding to the predicted trajectory is determined.

[0193] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0194] Fig.10FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Fig.10 As shown, the electronic device includes but is not limited to: a processor 1001 and a memory 1002 .

[0195] The memory 1002 is used to store executable instructions of the processor 1001. It can be understood that the processor 1001 is configured to execute instructions to implement the image acquisition method in the above embodiment.

[0196] It should be noted that those skilled in the art can understand that Fig.10 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device may include Fig.10 More or fewer components may be shown, or certain components may be combined, or the components may be arranged differently.

[0197] The processor 1001 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 1002, and calling data stored in the memory 1002, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 1001 may include one or more processing units. Optionally, the processor 1001 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1001.

[0198] The memory 1002 may be used to store software programs and various data. The memory 1002 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application program required by at least one functional module (such as a determination unit, a processing unit, etc.), etc. In addition, the memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0199] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1002 including instructions. The above instructions can be executed by a processor 1001 of an electronic device to implement the method in the above embodiment.

[0200] In actual implementation, Fig. 9 The functions of the processing unit 901 and the acquisition unit 902 in Fig.10The processor 1001 in the embodiment calls the computer program stored in the memory 1002 to implement. The specific execution process can refer to the description of the method part in the above embodiment, which will not be repeated here.

[0201] Optionally, the computer-readable storage medium may be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0202] In an exemplary embodiment, the present application also provides a computer program product including one or more instructions, and the one or more instructions can be executed by the processor 1001 of the electronic device to complete the method in the above embodiment.

[0203] It should be noted that when the instructions in the above-mentioned computer-readable storage medium or one or more instructions in the computer program product are executed by the processor of the electronic device, the various processes of the above-mentioned method embodiment are implemented, and the same technical effect as the above-mentioned method can be achieved. To avoid repetition, they will not be repeated here.

[0204] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0205] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0206] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0207] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0208] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or part of the prior art that contributes to the prior art or part or all of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks or optical disks.

[0209] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto, and any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A trajectory prediction method, characterized in that: The method comprises: Acquire a first position sequence feature and a lane segment sequence feature of the vehicle to be predicted; the first position sequence feature is used to characterize the position feature of the vehicle to be predicted at multiple historical moments; the lane segment sequence feature is used to characterize the position relationship between the vehicle to be predicted and the surrounding lane segments; Performing an attention mechanism operation based on the first position sequence feature and the lane segment sequence feature to obtain a first lane segment fusion feature; Based on the first lane segment fusion feature, a predicted trajectory of the to-be-predicted vehicle is obtained.

2. The method according to claim 1, characterized in that The obtaining the predicted trajectory of the to-be-predicted vehicle based on the first lane segment fusion feature includes: Acquire vehicle interaction features and second lane segment fusion features corresponding to surrounding vehicles of the vehicle to be predicted; the vehicle interaction features are used to characterize the positional relationship between the vehicle to be predicted and the surrounding vehicles; Performing an attention mechanism operation based on the first lane segment fusion feature, the vehicle interaction feature, and the second lane segment fusion feature to obtain a target fusion feature; A predicted trajectory of the vehicle to be predicted is generated according to the target fusion features.

3. The method according to claim 1, characterized in that The step of obtaining the first position sequence feature of the vehicle to be predicted includes: Obtaining historical position information of the vehicle to be predicted; Processing the historical position information based on a first multi-layer perceptron MLP to obtain a second position sequence feature of the vehicle to be predicted; Fusing the second position sequence feature and the time embedding parameter to obtain a third position sequence feature; An attention mechanism operation is performed on the third position sequence feature to obtain the first position sequence feature.

4. The method according to claim 1 or 3, characterized in that: Obtain the lane segment sequence features of the vehicle to be predicted, including: Obtaining the surrounding lane segment position information of the vehicle to be predicted; The surrounding lane segment position information is processed based on a second multi-layer perceptron MLP to obtain the lane segment sequence features.

5. The method according to claim 1, characterized in that The performing an attention mechanism operation based on the first position sequence feature and the lane segment sequence feature to obtain a first lane segment fusion feature includes: Mapping the first position sequence feature into a first query vector based on a third multi-layer perceptron MLP; Mapping the first splicing feature into a first key vector and a first value vector based on a fourth multi-layer perceptron MLP; the first splicing feature is a feature obtained by splicing the first position sequence feature and the lane segment sequence feature; The attention mechanism operation is performed based on the first query vector, the first key vector, and the first value vector to obtain the first lane segment fusion feature.

6. The method according to claim 2, characterized in that The obtaining of vehicle interaction features includes: Obtaining the position information of surrounding vehicles of the vehicle to be predicted; The surrounding vehicle position information is processed based on a fifth multi-layer perceptron MLP to obtain the vehicle interaction feature.

7. The method according to claim 2 or 6, characterized in that: The performing an attention mechanism operation based on the first lane segment fusion feature, the vehicle interaction feature, and the second lane segment fusion feature to obtain a target fusion feature includes: Mapping the first lane interaction feature into a second query vector based on a sixth multi-layer perceptron MLP; Mapping the second splicing feature into a second key vector and a second value vector based on the seventh multi-layer perceptron MLP; the second splicing feature is a feature obtained by splicing the first lane segment fusion feature, the vehicle interaction feature and the second lane segment fusion feature; The attention mechanism operation is performed based on the second query vector, the second key vector and the second value vector to obtain the target fusion feature.

8. A method for training a trajectory prediction model, characterized in that: The method comprises: Acquire a first position sequence feature and a lane segment sequence feature of the vehicle at a first sample moment; the first position sequence feature represents the position of the vehicle at multiple historical moments before the first sample moment; the lane segment sequence feature represents the position relationship between the vehicle and the surrounding lane segments at the multiple historical moments; Obtaining a target lane segment true value label and a true trajectory of the vehicle corresponding to the second sample time; The first position sequence feature and the lane segment sequence feature are processed based on the trajectory prediction model to obtain a predicted lane segment label and a predicted trajectory corresponding to the second sample moment; the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence feature and the lane segment feature to obtain a first lane segment fusion feature; based on the first lane segment fusion feature, the predicted trajectory is obtained, and a predicted lane segment label corresponding to the predicted trajectory is determined; Calculating a first loss based on the target lane segment true value label and the predicted lane segment label; Calculating a second loss based on the true trajectory and the predicted trajectory; The trajectory prediction model is adjusted according to the first loss and the second loss.

9. The method according to claim 8, characterized in that The first position sequence features and lane segment sequence features of the vehicle at the first sample time are obtained, including: Acquire historical position information of the vehicle at the multiple historical moments and surrounding lane segment position information of the to-be-predicted vehicle at the multiple historical moments; The historical position information and the surrounding lane segment position information are processed based on the trajectory prediction model to obtain the first position sequence feature and the lane segment sequence feature respectively.

10. The method according to claim 8, characterized in that The method further comprises: Acquire a vehicle interaction feature of the vehicle and a second lane segment fusion feature corresponding to surrounding vehicles of the vehicle; the vehicle interaction feature is used to characterize the positional relationship between the vehicle and the surrounding vehicles at the first sample time; The first position sequence feature and the lane segment sequence feature are processed based on the trajectory prediction model to obtain a predicted lane segment label and a predicted trajectory corresponding to the second sample time, including: The first lane segment fusion feature, the vehicle interaction feature and the second lane segment fusion feature are processed based on the trajectory prediction model to obtain a predicted lane segment label and a predicted trajectory corresponding to the second sample moment; the trajectory prediction model is used to perform an attention mechanism operation based on the first position sequence feature and the lane segment feature to obtain a first lane segment fusion feature; an attention mechanism operation is performed based on the first lane segment fusion feature, the vehicle interaction feature and the second lane segment fusion feature to obtain a target fusion feature; a predicted trajectory of the vehicle is generated according to the target fusion feature, and a predicted lane segment label corresponding to the predicted trajectory is determined.

11. A trajectory prediction device, characterized in that: The device comprises: a processing unit and an acquisition unit.

12. A vehicle, characterized in that: The vehicle comprises the device of claim 11.

13. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7 or to implement the method according to any one of claims 8 to 10.

14. A computer-readable storage medium, characterized in that: When the computer-executable instructions stored in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can perform the method according to any one of claims 1 to 7 or implement the method according to any one of claims 8 to 10.

15. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer, the computer is enabled to execute the method according to any one of claims 1 to 7 or to implement the method according to any one of claims 8 to 10.

Citation Information

Cited By

  • Trajectory prediction method and apparatus, and vehicle

    WO2026157298A1