Vehicle trajectory planning method and device, vehicle and storage medium
By fusing trajectory, map, and image information through encoding and decoding processes, the accuracy of automatic driving trajectory planning is improved, addressing the limitations of existing technologies that rely on limited perception data.
Patent Information
- Application Number
- CN202410053412.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-15
AI Technical Summary
In existing autonomous driving technologies, trajectory planning depends on less perceived information and is not fully utilized, resulting in poor accuracy of trajectory planning results.
By fusing the image of the target vehicle's surrounding environment, the object's historical trajectory information and road map, trajectory, map and image features are extracted and encoded, and the cross attention mechanism is used to decode, and trajectory prediction is performed to improve planning accuracy.
Improve the accuracy and reliability of trajectory planning, enhance the predictive ability of future trajectories, avoid collisions and comply with traffic rules.
Smart Images

Figure CN120313591A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular, to a vehicle trajectory planning method, apparatus, vehicle, and storage medium. Background Art
[0002] In recent years, the development of automotive electrification, intelligence, and networking has greatly promoted automotive assisted driving and autonomous driving technologies at all levels.
[0003] When performing autonomous driving, it is necessary to control the vehicle driving based on the result of trajectory planning. In related technologies, trajectory planning is performed based on the acquired perception information. However, the perception information relied on for trajectory planning is less, and the relationship between perception information is not fully utilized, resulting in poor accuracy of the trajectory planning result. Summary of the Invention
[0004] The present application aims to solve at least one of the technical problems in the related technologies to some extent.
[0005] To this end, the present application proposes a vehicle trajectory planning method, apparatus, vehicle, and storage medium. By fusing and encoding, the association relationship between trajectory information, map information, and image information is extracted. Furthermore, trajectory prediction is performed based on various position information and fusion features related to the target vehicle, improving the accuracy of the trajectory planning of the target vehicle.
[0006] An embodiment of one aspect of the present application proposes a vehicle trajectory planning method, including:
[0007] Obtaining an image of the environment around the target vehicle, historical trajectory information of an object in the surrounding environment, and a road map;
[0008] Fusing and encoding the historical trajectory features of the object, the map features of the road map, and the image features of the image to obtain a fusion feature;
[0009] Decoding according to the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion feature to obtain a target decoding feature;
[0010] Performing trajectory prediction according to the target decoding feature to obtain the trajectory planning result of the target vehicle.
[0011] An embodiment of another aspect of the present application proposes a vehicle trajectory planning apparatus, including:
[0012] An obtaining module, configured to execute obtaining an image of the environment around the target vehicle, historical trajectory information of an object in the surrounding environment, and a road map;
[0013] An encoding module, configured to perform fusion encoding on the historical trajectory features of the object, the map features of the road map, and the image features of the image to obtain fusion features;
[0014] A decoding module, configured to perform decoding according to the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features to obtain target decoding features;
[0015] A prediction module, configured to perform trajectory prediction according to the target decoding features to obtain the trajectory planning result of the target vehicle.
[0016] Another embodiment of this application provides a vehicle, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing aspect is implemented.
[0017] Another embodiment of this application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing aspect is implemented.
[0018] Another embodiment of this application provides a computer program product, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing aspect is implemented.
[0019] The vehicle trajectory planning method, device, vehicle, and storage medium provided by this application acquire an image of the surrounding environment of the target vehicle, historical trajectory information of an object in the surrounding environment, and a road map, perform fusion encoding on the historical trajectory features of the object, the map features of the road map, and the image features of the image to obtain fusion features, perform decoding according to the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features to obtain target decoding features, and perform trajectory prediction according to the target decoding features to obtain the trajectory planning result of the target vehicle. Thus, by performing fusion encoding based on trajectory features, map features, and image features, the fusion encoding extracts the association relationships between different types of features. Therefore, trajectory prediction is performed based on the fusion encoding, improving the accuracy of trajectory prediction.
[0020] The additional aspects and advantages of this application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0022] Figure 1 Schematic flowchart of a vehicle trajectory planning method provided by an embodiment of the present application;
[0023] Figure 2 Schematic flowchart of another vehicle trajectory planning method provided by an embodiment of the present application;
[0024] Figure 3 Schematic flowchart of another vehicle trajectory planning method provided by an embodiment of the present application;
[0025] Figure 4 Schematic structural diagram of a vehicle trajectory planning device provided by an embodiment of the present application;
[0026] Figure 5 Schematic structural diagram of a vehicle provided by an embodiment of the present application. Detailed implementation manners
[0027] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.
[0028] The vehicle trajectory planning method, device, vehicle, and storage medium of the embodiments of the present application will be described below with reference to the accompanying drawings.
[0029] Figure 1 Schematic flowchart of a vehicle trajectory planning method provided by an embodiment of the present application.
[0030] The execution subject of the vehicle trajectory planning method of the embodiments of the present application is a vehicle trajectory planning device, which can be set in a vehicle-mounted device, and is not limited in this embodiment.
[0031] As Figure 1 shown, the method may include the following steps:
[0032] Step 101, obtain an image of the surrounding environment of the target vehicle, historical trajectory information of objects in the surrounding environment, and a road map.
[0033] Among them, the target vehicle is the vehicle to be subjected to trajectory planning. The image of the surrounding environment of the target vehicle may be one image or multiple images, for example, multiple panoramic images collected by multiple panoramic cameras installed on the target vehicle.
[0034] Among them, objects in the surrounding environment may be obstacles during the driving of the target vehicle. As an implementation method, the images of the surrounding environment of the target vehicle can be subjected to object recognition to identify the objects included in the images, and the objects can be one or more. As another implementation method, the objects in the surrounding environment can be sensed by sensors installed on the target vehicle. For example, the objects in the surrounding environment are determined through point cloud data collected by radar sensors and the like.
[0035] Among them, the historical trajectory information of the object includes multiple trajectory points, and the state information of the object observable at each trajectory point includes position coordinates, speed, angle, etc. As an example, the number of objects is Na, and the historical trajectory information of each object is expressed as The historical trajectory information of Na objects can be expressed as Among them, t is the number of trajectory points in the historical trajectory or the number of historical frames, Ca is the number of observable states of an object in one frame, and the observable states include information such as position, speed, and angle.
[0036] Among them, the road map is usually represented by a road map vector. The road map includes map information. The road map is the map of the current driving area of the target vehicle. The road map can be divided into vector line segments according to the set position requirements. For example, the number of vector line segments in the road map is Nm, the number of points on each vector line segment is n, and the number of attributes of each point is Cm. Then the map vector information can be expressed as Among them, the attributes include position coordinates, line segment types, etc.
[0037] Step 102: Fuse and encode the historical trajectory features of the object, the map features of the road map, and the image features of the image to obtain fused features.
[0038] Among them, for feature extraction of the image, as an implementation method, first, the image is input into a spatial encoding network for feature extraction to obtain feature maps of multiple scales of the image. As an implementation method, for the image, Regnet800M in the image spatial encoding network is used as the image spatial encoder for feature extraction and encoding of multiple scales to generate feature maps of multiple scales of the image. Furthermore, the feature maps of multiple scales of the image are input into a multi-scale feature fusion network for feature fusion to obtain image features. As an implementation method, the BiFPN structure in the image spatial encoding network is used to perform feature fusion on the feature maps of different scales of the image, and finally a feature map with both low-level geometric image features and high-level semantic image features is obtained, and the image features are included in the feature map.
[0039] Among them, feature extraction is respectively performed on the historical trajectory information of the object and the road map to obtain the historical trajectory feature of the object and the map feature of the road map. The object is at least one. In the case where there are multiple objects, as an implementation manner, a multi-layer perceptron (MLP) and a max pooling layer are used to respectively perform feature extraction on the historical trajectories of each object among the at least one object to obtain the historical trajectory features of the at least one object. As an example, the at least one object is n, and the historical trajectory feature can be an n*m feature matrix. The feature matrix includes n columns, that is, the features of n objects. Among them, each column is the feature of one object, and each row is the feature dimension of one object. Similarly, feature encoding is performed on the road map to obtain the map feature.
[0040] In the embodiments of the present application, during the above-mentioned different information encoding processes, the encoding features of different information are independent of each other and have no association. In order to improve the accuracy of trajectory planning, during the process of performing trajectory planning on the target vehicle, on the one hand, it is necessary to avoid the target vehicle from colliding with surrounding objects, and on the other hand, it is necessary to comply with traffic rules. For example, it is necessary to drive forward along the lane line, the speed at the intersection should meet the speed limit requirements, and it is also necessary to comply with traffic light driving rules, etc. Therefore, in the embodiments of the present application, the historical trajectory feature indicating the trajectory of other vehicles, the map feature indicating traffic and road conditions, and the image feature indicating the overall environmental information are fused and encoded to obtain a fusion feature. That is to say, in the fusion feature, the historical trajectory feature, the map feature, and the image feature are mutually fused, and the features include each other, realizing that the information amount and expression ability carried by the obtained fusion feature are enhanced by fusing multiple features.
[0041] Step 103: Decode according to the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion feature to obtain the target decoding feature.
[0042] Among them, for the trajectory planning of the target vehicle, it needs to be determined based on the current positioning information of the target vehicle. The positioning information includes the current position coordinates of the target vehicle. The set position in the road map is the coordinate of the central position of multiple vector line segments divided in the road map. The positions of each pixel point in the image indicate the positions of the image features in the image, and the image features include the feature information of the object, such as the orientation and speed of the object.
[0043] In one implementation of the embodiment of the present application, based on the cross-attention mechanism, the current positioning information of the target vehicle, the positions of each pixel point in the environmental image, the set positions in the road map, and the fusion features are decoded, so that during the decoding process, the fusion features are decoded based on the current positioning information of the target vehicle, the positions of each pixel point in the environmental image, and the set positions in the road map, obtaining the target decoded features, such that the target decoded features fully consider the position information and fusion feature information related to the future trajectory of the target vehicle during decoding, improving the accuracy of the target decoded features.
[0044] Step 104, perform trajectory prediction according to the target decoded features to obtain the trajectory planning result of the target vehicle.
[0045] In one implementation of the embodiment of the present application, the target decoded features are used for trajectory prediction through a multi-layer perceptron (MLP) to obtain the trajectory planning result of the target vehicle, improving the accuracy of trajectory planning.
[0046] In the vehicle trajectory planning method of the embodiment of the present application, an image of the environment around the target vehicle, the historical trajectory information of the objects in the surrounding environment, and the road map are acquired. The historical trajectory features of the objects, the map features of the road map, and the image features of the image are fused and encoded to obtain fusion features. According to the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features, decoding is performed to obtain the target decoded features. Trajectory prediction is performed according to the target decoded features to obtain the trajectory planning result of the target vehicle. Thus, the correlation relationship between the trajectory information, the map information, and the image information is extracted through fusion encoding. Furthermore, trajectory prediction is performed based on various position information and fusion features related to the target vehicle, improving the accuracy of the trajectory planning of the target vehicle.
[0047] Based on the above embodiments, Figure 2 is a schematic flowchart of another vehicle trajectory planning method provided by the embodiment of the present application, as Figure 2 shown, this method includes the following steps:
[0048] Step 201, acquire an image of the environment around the target vehicle, the historical trajectory information of the objects in the surrounding environment, and the road map.
[0049] Among them, Step 201 can refer to the explanation in the foregoing embodiments, with the same principle and will not be elaborated here.
[0050] Step 202, perform feature merging on the historical trajectory features, map features, and image features to obtain the merged features.
[0051] In an implementation of the embodiment of the present application, the historical trajectory feature, the map feature, and the image feature are concatenated together in an hconcat manner to obtain a combined feature. Among them, when performing feature combination, the position information of the historical trajectory feature, the map feature, and the image feature in the combined feature is recorded in the combined feature. Each position information is used to split out the fused historical trajectory feature, the fused map feature, and the fused image feature after the subsequent fusion coding, that is, the first trajectory fusion feature, the first map fusion feature, and the first image fusion feature.
[0052] That is, the combined feature I1 = hconcat(A, M, P); where A is the historical trajectory feature, M is the map feature, and P is the image feature.
[0053] Step 203, input the combined feature into the fusion coding of the self-attention encoding module of the trained prediction model to obtain an intermediate feature.
[0054] Among them, the prediction model is trained based on training samples. As an implementation, training samples are obtained. The training samples include the first image of the environment around the target vehicle, the first historical trajectory information of the objects in the surrounding environment, and the first road map. Input the first image, the first historical trajectory information, and the first road map into the prediction model to obtain the predicted trajectory planning result output by the prediction model. Based on the difference between the first trajectory planning result and the ground truth carried by the training sample, that is, the true trajectory planning result, determine the loss function, and adjust the parameters of the prediction model based on the loss function to train the prediction model. In response to the loss function being less than the set threshold, or the change rate of the loss function being less than the threshold after the model iterates a set number of times, it is considered that the training of the prediction model is completed.
[0055] In the embodiment of the present application, the self-attention encoding module can fully fuse these three types of information to achieve mutual inclusion. That is, the encoded result of the historical trajectory information incorporates the map information and the image information, the encoded result of the map information incorporates the trajectory information and the image information, and the encoded result of the image information incorporates the trajectory information and the map information, realizing the mutual fusion of features, facilitating the reference to the correlation relationship between different information in the subsequent decoding process, and improving the accuracy of the obtained decoded features.
[0056] As an example, the intermediate feature I2 = FFN(MHSA(I1));
[0057] Among them, MHSA represents a multi-head self-attention encoder, which performs the encoding operation of multi-head self-attention. FFN represents a feed-forward network composed of multiple perceptrons, and the input of the feed-forward network is the output of the self-attention encoder.
[0058] Step 204: Split the features according to the position information of the historical trajectory feature, map feature, and image feature in the intermediate features to obtain the first trajectory fusion feature of the object, the first map fusion feature of the road map, and the first image fusion feature of the image.
[0059] In an implementation manner of the embodiment of the present application, the splitting function Split is used to split according to the position information of the historical trajectory feature, map feature, and image feature in the intermediate features to obtain the first trajectory fusion feature of the object, the first map fusion feature of the road map, and the first image fusion feature of the image.
[0060] It should be understood that when driving a vehicle, it is necessary to continuously observe and remember the trajectory information of other vehicles around, the road structure and traffic sign information (i.e., map information), and the overall environmental information (i.e., image information). These three types of information are not isolated but are interrelated. Therefore, we combine these three types of information in the hconcat manner and send them into the self-attention encoding module. The self-attention encoding module can fully fuse these three types of information to achieve mutual inclusion. In the final output result, the encoded result of the trajectory information, that is, the first trajectory fusion feature, incorporates the map information and the image information. The encoded result of the map information, that is, the first map fusion feature, incorporates the trajectory information and the image information. The encoded result of the image information, that is, the first image fusion feature, incorporates the trajectory information and the map information, making the output various fusion features have stronger expression capabilities.
[0061] Step 205: Decode according to the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features to obtain the target decoded features.
[0062] In the embodiment of the present application, future trajectory prediction is performed based on the first trajectory fusion feature of the object to obtain the future trajectory of the object. As an implementation manner, the direct linear regression method can be used, that is, future trajectory prediction is performed through the multi-layer perceptron MLP to obtain the future trajectory of the object. Among them, the future trajectory is relative to the historical trajectory. For example, if the historical trajectory is the trajectory in the t1 - t2 time period, then the future trajectory is the trajectory in the t2 - t3 time period.
[0063] As an example, the future trajectory T2 is determined by the following method:
[0064] T2 = MLP(A 融合1 );
[0065] Among them, the multi-layer perceptron MLP is used to map a set of input vectors to a set of output vectors, and A 融合1 is the first trajectory fusion feature.
[0066] Furthermore, encode the future trajectory to obtain future trajectory features. As an implementation, the future trajectory feature T1 is achieved through the following method:
[0067] T1 = maxpool(MLP(T2));
[0068] That is, input the future trajectory T2 into the multi-layer perceptron MLP for processing to obtain intermediate features, and perform max pooling operation (maxpool) on the intermediate features to obtain future trajectory features. Among them, the multi-layer perceptron can include multiple processing layers, and each processing layer includes a fully connected layer and a pooling layer.
[0069] Furthermore, splice the future trajectory features and the first trajectory fusion features to obtain the second trajectory fusion features of the object. Therefore, the second trajectory fusion features include both past historical trajectory information and future trajectory information, and can be used as important reference information for obstacles in the trajectory planning of the target vehicle.
[0070] Finally, input the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, the second trajectory fusion features, the first map fusion features, and the first image fusion features into the decoding module of the prediction model to decode and obtain the target decoding features. Among them, the specific decoding algorithm will be explained in subsequent embodiments and will not be elaborated here.
[0071] Step 206, perform trajectory prediction according to the target decoding features to obtain the trajectory planning result of the target vehicle.
[0072] Among them, the explanations in the foregoing embodiments are also applicable to this embodiment, with the same principle and will not be elaborated here.
[0073] In the vehicle trajectory planning method of the embodiment of the present application, after merging the historical trajectory features, map features, and image features, input them into the self-attention encoding module of the trained prediction model for fusion encoding to obtain intermediate features. According to the position information of the historical trajectory features, map features, and image features in the intermediate features, perform feature splitting to obtain the first trajectory fusion features of the object, the first map fusion features of the road map, and the first image fusion features of the image. Through the self-attention encoding mechanism, fuse and encode the historical trajectory features, map features, and image features, so that the first trajectory fusion features obtained by encoding fuse image features and map features, the first map fusion features fuse trajectory features and image features, and the first image fusion features fuse trajectory features and map features. Furthermore, based on the first trajectory fusion features, the first map fusion features, and the first image fusion features, perform decoding to obtain the target decoding features, which can improve the accuracy and reliability of the decoding feature determination result, and further improve the accuracy and reliability of the trajectory planning result.
[0074] Based on the above embodiments, an embodiment of the present application provides another vehicle trajectory planning method. Figure 3 As shown in the flowchart of another vehicle trajectory planning method provided by the embodiment of the present application, Figure 3 the method includes the following steps:
[0075] Step 301, obtain an image of the surrounding environment of the target vehicle, historical trajectory information of objects in the surrounding environment, and a road map.
[0076] Step 302, fuse and encode the historical trajectory features of the object, the map features of the road map, and the image features of the image to obtain fused features.
[0077] Among them, the fused features include: the first trajectory fused feature of the object, the first map fused feature of the road map, and the first image fused feature of the image.
[0078] Step 303, perform future trajectory prediction based on the first trajectory fused feature of the object to obtain the future trajectory of the object, and encode the future trajectory to obtain future trajectory features.
[0079] Step 304, splice the future trajectory features and the first trajectory fused feature to obtain the second trajectory fused feature of the object.
[0080] As an implementation manner, after performing a vconcat operation on the first trajectory fused feature of the object and the future trajectory features, a second trajectory fused feature including past and future trajectory information is obtained.
[0081] Among them, Steps 301 - 304 can refer to the explanations in the foregoing embodiments, with the same principle, which will not be elaborated here.
[0082] Step 305, obtain a set first query vector corresponding to the trajectory features, a set second query vector corresponding to the image features, and a set third query vector corresponding to the map features.
[0083] In an implementation manner of the embodiment of the present application, the prediction model includes an encoding module and a decoding module, and the decoding model includes three sub - decoding modules. The three sub - decoding modules process in parallel, namely the first sub - decoding module for decoding trajectory information, the second sub - decoding module for decoding map information, and the third sub - decoding module for decoding image information.
[0084] Among them, the first sub-decoding module includes a set first query vector. The first query vector is the query parameter corresponding to the first sub-decoding module in the trained prediction model, that is, q1 (query), and this query parameter is used to query trajectory information. The second sub-decoding module includes a set second query vector. The second query vector is the query parameter corresponding to the second sub-decoding module in the trained prediction model, that is, q2, and this query parameter is used to query map information. The third sub-decoding module includes a set third query vector. The third query vector is the query parameter corresponding to the third sub-decoding module in the trained prediction model, that is, q3, and this query parameter is used to query image information.
[0085] Step 306: Input the first query vector, the second trajectory fusion feature, the current positioning information of the target vehicle, and the end position in the second trajectory into the first sub-decoding module in the decoding module to decode and obtain a first decoding vector.
[0086] In the embodiment of this application, according to the current positioning information of the target vehicle and the first query vector, a first target query vector is determined. As an implementation, the current positioning information of the target vehicle is position-encoded to obtain a target vehicle position-encoding feature, and the target vehicle position-encoding feature and the first query vector are added to obtain the first target query vector, that is, q1' = q1 + PE(pos ego ), where pos ego is the current positioning information of the target vehicle, PE represents position-encoding the two-dimensional coordinates to obtain the vehicle position-encoding feature, q1 is the first query vector, and the first target query vector is q1'.
[0087] According to the second trajectory fusion feature and the end position of the second trajectory, a third trajectory fusion feature is determined. As an implementation, the end position of the second trajectory is position-encoded to obtain an object position-encoding feature, and the object position-encoding feature and the second trajectory fusion feature are added to obtain the third trajectory fusion feature. That is, A 融合3 = A 融合2 + PE(pos obj ); where PE(pos obj ) is the object position-encoding feature.
[0088] Furthermore, input the first target query vector, the third trajectory fusion feature, and the second trajectory fusion feature into the first sub-decoding module for decoding to obtain the first decoding vector. As an implementation, the first sub-decoding module is a cross-attention decoder, which decodes the first target query vector, the third trajectory fusion feature, and the second trajectory fusion feature based on the cross-attention algorithm, that is, performs an inner product of the first target query vector and the third trajectory fusion feature and normalizes it to obtain attention weights, and weights the second trajectory fusion feature according to the attention weights to obtain the first decoding vector.
[0089] As an example, the inputs to the first sub-decoding module are the first target query vector, the third trajectory fusion feature, and the second trajectory fusion feature. Among them, the first target query vector is the first parameter q (query) of the cross-attention decoder, the third trajectory fusion feature is the second parameter k (key) of the cross-attention decoder, and the second trajectory fusion feature is the third parameter v (value) of the cross-attention decoder.
[0090] The first decoding vector q A is determined by the following method:
[0091] q A = MHCA(query = q1 + PE(pos ego ), key = A 融合2 + PE(pos ego ), value = A 融合2 );
[0092] Among them, MHCA represents the decoding operation of the multi-head cross-attention decoder.
[0093] Step 307: Input the second query vector, the first map fusion feature, the current positioning information of the target vehicle, and the set position in the road map into the second sub-decoding module in the decoding module for decoding to obtain the second decoding vector.
[0094] In the embodiment of the present application, according to the current positioning information of the target vehicle and the second query vector, the second target query vector is determined. As an implementation, the current positioning information of the target vehicle is position-encoded to obtain the target vehicle position encoding feature, and the target vehicle position encoding feature and the second query vector are added to obtain the second target query vector, that is, q2' = q2 + PE(pos ego ).
[0095] Determine the second map fusion feature based on the first map fusion feature and the set position in the road map. As an implementation, perform position encoding on the set position in the road map to obtain the encoded feature of the set position in the map, and add the encoded feature of the set position in the map and the first map fusion feature to obtain the second map fusion feature, that is, M 融合2 = M 融合1 + PE(pos map ); where PE(pos map ) is the encoded feature of the set position in the map.
[0096] Furthermore, input the second query target, the second map fusion feature, and the first map fusion feature into the second sub-decoding module for decoding to obtain the second decoding vector. As an implementation, the second sub-decoding module is a cross-attention decoder, and decode the second target query vector, the second map fusion feature, and the first map fusion feature based on the cross-attention algorithm, that is, take the inner product of the second target query vector and the second map fusion feature and normalize it to obtain the attention weights, and weight the first map fusion feature according to the attention weights to obtain the second decoding vector.
[0097] As an example, the inputs to the second sub-decoding module are the second target query vector, the second map fusion feature, and the first map fusion feature. Among them, the second target query vector is the first parameter q (query) of the cross-attention decoder, the second map fusion feature is the second parameter k (key) of the cross-attention decoder, and the first map fusion feature is the third parameter v (value) of the cross-attention decoder.
[0098] The second decoding vector q M is determined in the following way:
[0099] q M = MHCA(query = q2 + PE(pos ego ), key = M 融合1 + PE(pos map ), value = M 融合1 ).
[0100] Step 308, input the third query vector, the first image fusion feature, the current positioning information of the target vehicle, and the positions of each pixel point in the image into the third sub-decoding module in the decoding module for decoding to obtain the third decoding vector.
[0101] In the embodiments of the present application, a third target query vector is determined according to the current positioning information of the target vehicle and the third query vector. As an implementation, the current positioning information of the target vehicle is position-encoded to obtain a target vehicle position encoding feature, and the target vehicle position encoding feature and the third query vector are added to obtain the third target query vector, that is, q3' = q3 + PE(pos ego ).
[0102] According to the first image fusion feature and the positions of each pixel point in the image, a second image fusion feature is determined. As an implementation, the positions of each pixel point in the image are position-encoded to obtain the encoding features of the positions of each pixel point in the image, and the encoding features of the positions of each pixel point in the image and the first image fusion feature are added to obtain the second image fusion feature, that is, P 融合2 = P 融合1 + PE(pos pixel ); where PE(pos pixel ) is the encoding feature of the positions of each pixel point in the image.
[0103] Furthermore, the third target query vector, the first image fusion feature, and the second image fusion feature are input into the third sub-decoding module for decoding to obtain an image query vector. As an implementation, the third sub-decoding module is a cross-attention decoder, and the third target query vector, the second image fusion feature, and the first image fusion feature are decoded based on the cross-attention algorithm, that is, the inner product of the third target query vector and the second image fusion feature is normalized to obtain attention weights, and the first image fusion feature is weighted according to the attention weights to obtain a third decoded vector.
[0104] As an example, the inputs of the third sub-decoding module are the third target query vector, the second image fusion feature, and the first image fusion feature, where the third target query vector is the first parameter q (query) of the cross-attention decoder, the second image fusion feature is the second parameter k (key) of the cross-attention decoder, and the first image fusion feature is the third parameter v (value) of the cross-attention decoder.
[0105] The third decoded vector q p is determined by the following method:
[0106] q p = MHCA(query = q3 + PE(pos ego ), key = P 融合1 + PE(pos pixel ), value = P 融合1 )
[0107] Step 309: Determine the target decoding feature according to the concatenation result of the first decoding vector, the second decoding vector, and the third decoding vector.
[0108] In an implementation of the embodiments of the present application, the first decoding vector, the second decoding vector, and the third decoding vector are concatenated to obtain a vector, which is processed by a multi-layer perceptron (MLP) to obtain the target decoding feature, enhancing the feature expression ability of the target decoding feature. Among them, the target decoding feature is determined by the following formula:
[0109] q 目标 = MLP(concat(q A , q M , q p ));
[0110] Step 310: Perform trajectory prediction according to the target decoding feature to obtain the trajectory planning result of the target vehicle.
[0111] In an implementation of the embodiments of the present application, the target decoding feature is input into the prediction module of the prediction model to obtain the trajectory planning result of the target vehicle. Among them, the prediction module is the prediction head for regressing the trajectory. Compared with the related technology that first generates a trajectory set based on sampling and then selects the trajectory that meets the optimization goal from the trajectory set as the trajectory planning result of the target vehicle, in the present application, the prediction result is directly obtained based on the target decoding feature, simplifying the prediction process and improving the prediction efficiency.
[0112] As an example, the trajectory planning result Traj of the target vehicle is:
[0113] Traj = RegHead(q 目标 );
[0114] Among them, RegHead represents the prediction head of the prediction module for regressing the trajectory, and Traj ∈ R T×5 , where T represents the length of the trajectory planning, and 5 represents 5 attributes, namely the x and y coordinates, the speeds in the x and y directions, and the angle, for a total of 5 attribute values.
[0115] In the vehicle trajectory planning method of the embodiments of the present application, during the decoding process according to the first trajectory fusion feature, the first image fusion feature, and the first map fusion feature output by the encoder, cross-attention is used for decoding to obtain the decoding features corresponding to the respective fusion features, the respective decoding features are concatenated to obtain the target decoding feature, and trajectory prediction is performed based on the target decoding feature to obtain the trajectory prediction result of the target vehicle. By performing decoding based on cross-attention, the accuracy of the decoded information is improved, and thus the accuracy of the trajectory prediction of the target vehicle is improved.
[0116] To implement the above embodiments, an embodiment of the present application further provides a vehicle trajectory planning device.
[0117] Figure 4 It is a schematic structural diagram of a vehicle trajectory planning device provided by an embodiment of the present application.
[0118] As Figure 4 shown, the device may include:
[0119] An acquisition module 41, configured to acquire an image of the environment around the target vehicle, historical trajectory information of objects in the surrounding environment, and a road map.
[0120] An encoding module 42, configured to perform fusion encoding according to the historical trajectory features of the objects, the map features of the road map, and the image features of the image to obtain fusion features.
[0121] A decoding module 43, configured to decode the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features to obtain target decoding features.
[0122] A planning module 44, configured to perform trajectory prediction according to the target decoding features to obtain a trajectory planning result of the target vehicle.
[0123] Furthermore, in an implementation manner of an embodiment of the present application, the encoding module 42 is further configured to perform:
[0124] Feature merging of the historical trajectory features, the map features, and the image features to obtain merged features; wherein, the merged features carry the position information of the historical trajectory features, the map features, and the image features;
[0125] Inputting the merged features into the encoding module of the self-attention of the trained prediction model for fusion encoding to obtain intermediate features;
[0126] Feature splitting according to the position information of the historical trajectory features, the map features, and the image features in the intermediate features to obtain a first trajectory fusion feature of the object, a first map fusion feature of the road map, and a first image fusion feature of the image.
[0127] In an implementation manner of an embodiment of the present application, the decoding module 43 is further configured to perform:
[0128] Performing future trajectory prediction based on the first trajectory fusion feature of the object to obtain the future trajectory of the object;
[0129] Encoding the future trajectory to obtain future trajectory features;
[0130] Concatenate the future trajectory feature and the first trajectory fusion feature to obtain the second trajectory fusion feature of the object;
[0131] Input the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, the second trajectory fusion feature, the first map fusion feature, and the first image fusion feature into the decoding module of the prediction model to decode and obtain the target decoding feature.
[0132] In an implementation manner of the embodiment of the present application, the decoding module 43 is further configured to execute:
[0133] Obtain a set first query vector corresponding to the trajectory feature, a set second query vector corresponding to the image feature, and a set third query vector corresponding to the map feature;
[0134] Input the first query vector, the second trajectory fusion feature, the current positioning information of the target vehicle, and the end position in the second trajectory into the first sub-decoding module in the decoding module to decode and obtain the first decoding vector;
[0135] Input the second query vector, the first map fusion feature, the current positioning information of the target vehicle, and the set positions in the road map into the second sub-decoding module in the decoding module to decode and obtain the second decoding vector;
[0136] Input the third query vector, the first image fusion feature, the current positioning information of the target vehicle, and the positions of each pixel point in the image into the third sub-decoding module in the decoding module to decode and obtain the third decoding vector;
[0137] Determine the target decoding feature according to the concatenation result of the first decoding vector, the second decoding vector, and the third decoding vector.
[0138] In an implementation manner of the embodiment of the present application, the decoding module 43 is further configured to execute:
[0139] Determine a first target query vector according to the current positioning information of the target vehicle and the first query vector;
[0140] Determine a third trajectory fusion feature according to the second trajectory fusion feature and the end position of the second trajectory;
[0141] Input the first target query vector, the third trajectory fusion feature, and the second trajectory fusion feature into the first sub-decoding module to decode and obtain the first decoding vector.
[0142] In an implementation manner of the embodiment of the present application, the decoding module 43 is further configured to execute:
[0143] Determine a second target query vector according to the current positioning information of the target vehicle and the second query vector;
[0144] Determine a second map fusion feature according to the first map fusion feature and the set position in the road map;
[0145] Input the second query target, the second map fusion feature, and the first map fusion feature into the second sub-decoding module for decoding to obtain the map decoding vector.
[0146] In an implementation manner of the embodiment of the present application, the decoding module 43 is further configured to execute:
[0147] Determine a third target query vector according to the current positioning information of the target vehicle and the third query vector;
[0148] Determine a second image fusion feature according to the first image fusion feature and the positions of each pixel point in the image;
[0149] Input the third target query vector, the first image fusion feature, and the second image fusion feature into the third sub-decoding module for decoding to obtain the third decoding vector.
[0150] It should be noted that the foregoing explanation of the method embodiment is also applicable to the device of this embodiment, and will not be repeated here.
[0151] In the vehicle trajectory planning device of the embodiment of the present application, an image of the surrounding environment of the target vehicle, historical trajectory information of objects in the surrounding environment, and a road map are obtained, and fusion coding is performed according to the historical trajectory features of the objects, the map features of the road map, and the image features of the image to obtain a fusion feature. Decoding is performed according to the current positioning information of the target vehicle, the positions of each pixel point in the image, the set position in the road map, and the fusion feature to obtain a target decoding feature, and trajectory prediction is performed according to the target decoding feature to obtain the trajectory planning result of the target vehicle. Thus, the correlation relationship between the trajectory information, map information, and image information is extracted through fusion coding. Furthermore, trajectory prediction is performed based on various position information and fusion features related to the target vehicle, improving the accuracy of the trajectory planning of the target vehicle.
[0152] To implement the above embodiment, the present application also proposes a vehicle, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing method embodiment is implemented.
[0153] To implement the above embodiments, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the method as described in the foregoing method embodiments.
[0154] To implement the above embodiments, the present application also provides a computer program product having stored thereon a computer program, which when executed by a processor, implements the method as described in the foregoing method embodiments.
[0155] Figure 5 FIG. 7 is a schematic structural diagram of a vehicle shown according to an exemplary embodiment. For example, vehicle 600 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. Vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0156] Referring to Figure 5 , vehicle 600 may include various subsystems. For example, the infotainment system 610, the perception system 620, the decision control system 630, the drive system 640, and the computing platform 650. Among them, vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem of vehicle 600 and each component may be interconnected in a wired or wireless manner.
[0157] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, and a navigation system, etc.
[0158] The perception system 620 may include several sensors for sensing information about the environment around vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), a lidar, a millimeter wave radar, an ultrasonic radar, and a camera device.
[0159] The decision control system 630 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0160] The drive system 640 may include components that provide motive power for vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a powertrain, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.
[0161] Some or all functions of vehicle 600 are controlled by computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652. The processor 651 may execute instructions 653 stored in the memory 652.
[0162] The processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0163] The memory 652 may be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0164] In addition to the instructions 653, the memory 652 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 652 can be used by the computing platform 650.
[0165] In the embodiments of the present disclosure, the processor 651 may execute the instructions 653 to complete all or part of the steps of the above method embodiments.
[0166] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0167] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0168] Any process or method description represented in a flowchart or described otherwise herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0169] The logic and / or steps represented in a flowchart or described otherwise herein, for example, may be considered as a sequenced list of executable instructions for implementing a logical function and may be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained, for example, electronically by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise appropriate processing if necessary, and then stored in a computer memory.
[0170] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0171] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0172] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0173] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A vehicle trajectory planning method, characterized in that, Including: Obtaining an image of the surrounding environment of the target vehicle, historical trajectory information of objects in the surrounding environment, and a road map; Performing fusion encoding based on the historical trajectory features of the objects, the map features of the road map, and the image features of the image to obtain fusion features; Performing decoding based on the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features to obtain target decoding features; Performing trajectory prediction based on the target decoding features to obtain the trajectory planning result of the target vehicle.
2. The method according to claim 1, wherein The performing fusion encoding based on the historical trajectory features of the objects, the map features of the road map, and the image features of the image to obtain fusion features includes: Performing feature merging on the historical trajectory features, the map features, and the image features to obtain merged features; wherein, the merged features carry the position information of the historical trajectory features, the map features, and the image features; Inputting the merged features into the encoding module of the self-attention of the trained prediction model for fusion encoding to obtain intermediate features; Performing feature splitting according to the position information of the historical trajectory features, the map features, and the image features in the intermediate features to obtain the first trajectory fusion feature of the object, the first map fusion feature of the road map, and the first image fusion feature of the image.
3. The method according to claim 2, wherein The performing decoding based on the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features to obtain target decoding features includes: Performing future trajectory prediction based on the first trajectory fusion feature of the object to obtain the future trajectory of the object; Encoding the future trajectory to obtain future trajectory features; Concatenating the future trajectory features and the first trajectory fusion feature to obtain the second trajectory fusion feature of the object; Inputting the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, the second trajectory fusion feature, the first map fusion feature, and the first image fusion feature into the decoding module of the prediction model for decoding to obtain target decoding features.
4. The method according to claim 3, characterized in that, The inputting the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, the second trajectory fusion feature, the first map fusion feature, and the first image fusion feature into the decoding module of the prediction model for decoding to obtain target decoding features includes: Obtaining a set first query vector corresponding to the trajectory features, a set second query vector corresponding to the image features, and a set third query vector corresponding to the map features; Inputting the first query vector, the second trajectory fusion feature, the current positioning information of the target vehicle, and the end position in the second trajectory into the first sub-decoding module in the decoding module for decoding to obtain a first decoding vector; Input the second query vector, the first map fusion feature, the current positioning information of the target vehicle, and the set position in the road map into the second sub-decoding module in the decoding module to decode and obtain a second decoded vector; Input the third query vector, the first image fusion feature, the current positioning information of the target vehicle, and the positions of each pixel point in the image into the third sub-decoding module in the decoding module to decode and obtain a third decoded vector; Determine the target decoded feature according to the concatenation result of the first decoded vector, the second decoded vector, and the third decoded vector.
5. The method according to claim 4, characterized in that The step of inputting the first query vector, the second trajectory fusion feature, the current positioning information of the target vehicle, and the end position in the second trajectory into the first sub-decoding module in the decoding module to decode and obtain a first decoded vector includes: Determine a first target query vector according to the current positioning information of the target vehicle and the first query vector; Determine a third trajectory fusion feature according to the second trajectory fusion feature and the end position of the second trajectory; Input the first target query vector, the third trajectory fusion feature, and the second trajectory fusion feature into the first sub-decoding module to decode and obtain a first decoded vector.
6. The method according to claim 4, wherein The step of inputting the second query vector, the first map fusion feature, the current positioning information of the target vehicle, and the set position in the road map into the second sub-decoding module in the decoding module to decode and obtain a second decoded vector includes: Determine a second target query vector according to the current positioning information of the target vehicle and the second query vector; Determine a second map fusion feature according to the first map fusion feature and the set position in the road map; Input the second query target, the second map fusion feature, and the first map fusion feature into the second sub-decoding module to decode and obtain the map decoded vector.
7. The method according to claim 4, characterized in that The step of inputting the third query vector, the first image fusion feature, the current positioning information of the target vehicle, and the positions of each pixel point in the image into the third sub-decoding module in the decoding module to decode and obtain a third decoded vector includes: Determine a third target query vector according to the current positioning information of the target vehicle and the third query vector; Determine a second image fusion feature according to the first image fusion feature and the positions of each pixel point in the image; Input the third target query vector, the first image fusion feature, and the second image fusion feature into the third sub-decoding module to decode and obtain the third decoded vector.
8. A vehicle trajectory planning device, characterized in that, It includes: An acquisition module, configured to acquire an image of the surrounding environment of the target vehicle, historical trajectory information of an object in the surrounding environment, and a road map; An encoding module, configured to perform fusion encoding according to the historical trajectory feature of the object, the map feature of the road map, and the image feature of the image to obtain a fusion feature; A decoding module, configured to perform decoding based on the current positioning information of the target vehicle, the positions of each pixel point in the image, the set positions in the road map, and the fusion features to obtain target decoding features; A prediction module, configured to perform trajectory prediction based on the target decoding features to obtain the trajectory planning result of the target vehicle.
9. A vehicle, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the mobile terminal, the mobile terminal is enabled to execute the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Trajectory prediction method and related device
CN121106344A